Social Screening API Coverage: Platforms, Signal Types and Depth

Social screening coverage is 3 questions vendors report as one: can you find the accounts, see the content, and understand it. Why no findings is ambiguous.

Ronak Shah
Growth at Phyllo
August 18, 2026
August 18, 2026
3D icons of a magnifying glass, a video play button and a globe
Summarize this article with AI
GeminiChatGPTClaudePerplexityGrok

Social screening coverage is 3 questions vendors report as one: can you find the accounts, see the content, and understand it. Why no findings is ambiguous.

This is some text inside of a div block.
  • Social screening coverage is 3 separate questions reported as one number: can you find the accounts, see the content, and understand it.
  • No findings carries 2 opposite meanings, the subject is clean or nothing was visible, and most reports never say which.
  • Video is the biggest gap, with one major vendor gating full video analysis to its Enterprise tier.
  • Language is the second gap, because a classifier tuned on English slurs misses the same content in Hindi, Tagalog or Arabic.
  • From 30 March 2026 US social media screening covers 15+ visa categories, making coverage a regulatory requirement.

Every screening vendor leads with a platform count, and it is the least informative number they could publish. Coverage in screening is not a field lookup. It is a chain of three dependent questions, and a failure at any link produces the same output as a clean subject: nothing found.

The three questions are reach, meaning can you find the subject's accounts at all; visibility, meaning can you see the content on the accounts you found; and comprehension, meaning can you actually classify what you can see. A tool can cover forty platforms and answer no on all three for a given subject.

What follows is each link in that chain, what breaks it, and how to express coverage as something a buyer can verify. If you want the definitional groundwork first, start with what social screening is.

Question one: can you find the accounts?

This is identity resolution and it is the hardest part of screening, not the easiest. A screening report on the wrong person is not a data quality problem, it is an adverse action against the wrong candidate.

Three things break reach, and none of them appear in a platform count.

Pseudonymous accounts

A subject's riskiest content is rarely on the account bearing their legal name. Pseudonymous accounts are frequently linkable, and a documented case makes the point: an F-1 visa applicant received a 221(g) refusal in July 2025 for failing to disclose a public, pseudonymous Reddit account that a search engine linked to his name. The account was findable. The screening question is whether your vendor found it.

Regional platforms

A tool covering Facebook, Instagram, X, LinkedIn, TikTok and YouTube looks comprehensive and can still miss most of a subject's footprint. The US DS-160 form carries a free-text field precisely because the dropdown does not cover everything, and guidance for Indian applicants names Sharechat, Koo, Moj and Josh among the platforms to declare.

So ask your vendor for the platform list, restricted to the countries your subject population actually comes from, rather than a global total. A vendor covering 40 global platforms and zero Indian regional platforms has zero coverage on a large share of an Indian applicant pool. That is not a small gap, and a count conceals it perfectly.

Name ambiguity at scale

Common names produce many candidate matches and the failure is asymmetric. A false negative means you miss a signal. A false positive means you attribute someone else's content to your subject and act on it. Ask what happens on an ambiguous match, whether a human reviews it, and whether the report states a match confidence rather than presenting a probable match as a certain one. This is what identity resolution and account linkage exist to solve.

Question two: can you see the content?

Finding an account and reading it are different achievements. Four things sit between the two, and all four are outside a vendor's control.

BarrierWhat it meansWhat a vendor can honestly do
Private accountsContent is not publicly visible. No compliant route existsReport the account as found but inaccessible. Nothing else
Deleted contentScreening sees current state. A subject can remove anything before a screenNote the review date. Public screening has no archive of what was removed
Ephemeral contentStories and similar formats expire within 24 hours and are never publicly archivedNothing, unless continuously monitoring from before the content posted
Lookback limitsVendors commonly offer 7 to 10 years of post history. Older content may be unreachable rather than absentState the window achieved per account, not the window offered

On private accounts, be clear about what is and is not on the table. Reviewing publicly visible content is the entire legitimate scope. Requesting credentials, or asking a subject to accept a connection request so private content becomes visible, runs into password protection laws in a large number of US states and is not something a vendor can solve with technology.

One context where this works differently: US visa screening. Applicants must disclose handles used over a five-year lookback including deleted accounts, and in many cases embassies instruct applicants to make accounts publicly viewable before the interview. The privacy barrier is addressed by disclosure obligation rather than by technical access, which is a legal mechanism, not a coverage capability.

Why is "no findings" the most dangerous line in a screening report?

Because it is produced by two opposite situations and almost no report distinguishes them. A subject with a decade of clean public activity and a subject whose accounts are all private both come back as nothing found.

Think about what that does to a hiring decision. In the first case, the absence of findings is evidence. In the second, it is the absence of evidence, which is a completely different input and arguably tells you the screen did not happen. Presented in identical wording, the two collapse into one signal that a reviewer will read as reassurance.

This is the single most important thing to ask a screening vendor about, and it is almost never on a feature list. The question is: does your report distinguish clean from invisible, and how?

A defensible report answers it structurally rather than in prose. It states how many accounts were found, how many were accessible, how many content items were reviewed, which modalities were analysed, and what lookback was actually achieved. Then "no findings" means something, because it comes with a denominator.

Question three: can you understand what you see?

This is where coverage claims quietly fall apart, because the content that matters most has moved to the formats that are hardest to analyse.

The four modalities

Serious content analysis needs all four. Most tools handle the first well, the second reasonably, and the last two poorly or not at all.

ModalityWhat it isTypical capabilityWhere it fails
Caption and comment textWritten content on and under a postStrong across vendorsCoded language, irony, in-group slang
On-screen textText burned into an image or video frameVariable. Needs OCR plus classificationStylised fonts, low contrast, moving text
Image contentWhat is depicted, including gestures and symbolsModerateContext, satire, and symbols that are regionally specific
Audio and videoSpoken content, and what happens across a clipWeak, and frequently a premium tierEverything. This is the expensive one

The commercial evidence is in the pricing. One established screening vendor lists full video analysis as an Enterprise-tier feature, above its per-report Professional tier. Capabilities get gated when they are expensive, and video analysis is expensive because it is genuinely hard.

Now put that next to where content lives. Short-form video is the dominant format on TikTok, Reels and Shorts, and a tool that reads only captions on those platforms is reporting on the description rather than the content. It will not find a threat spoken aloud in a fifteen-second clip with an innocuous caption. That is not an edge case, it is the median post on those platforms.

Language coverage

The second comprehension gap, and the one most likely to affect a global employer. A classifier tuned on English hate speech and English threat language will underperform on the same content in Hindi, Tagalog, Spanish, Arabic or Portuguese. Slurs are language-specific and often region-specific, coded language does not translate, and a model that has seen little training data in a language will produce more false negatives in it.

The consequence is a fairness problem as well as a coverage one. If your screen is more sensitive in English than in Tagalog, you are applying a different standard to different candidate populations, which is exactly the inconsistency that makes a screening process hard to defend.

Ask for the language list, ask which languages are natively classified versus machine translated first, and ask for measured performance per language rather than a claim of multilingual support.

What should a coverage claim actually look like?

A per-subject completeness score, computed and reported alongside the findings. This is buildable, it is the honest way to express screening coverage, and I have not seen a vendor publish it.

# Report coverage per subject, not per vendor.
# "No findings" is only meaningful next to this.
MODALITIES = ["text", "on_screen_text", "image", "audio_video"]
def completeness(subject):
    declared   = subject.declared_handles          # from the subject or the form
    found      = subject.accounts_found            # after identity resolution
    accessible = [a for a in found if a.public]    # private accounts excluded
    return {
        # REACH
        "handles_declared":     len(declared),
        "accounts_found":       len(found),
        "unmatched_declared":   len([h for h in declared
                                     if h not in {a.handle for a in found}]),
        "match_confidence_min": min((a.confidence for a in found), default=None),
        # VISIBILITY
        "accounts_accessible":  len(accessible),
        "accounts_private":     len(found) - len(accessible),
        "items_reviewed":       sum(a.items_reviewed for a in accessible),
        "lookback_achieved_yrs": min((a.oldest_item_age for a in accessible),
                                     default=0),
        # COMPREHENSION
        "modalities_analysed":  [m for m in MODALITIES
                                 if subject.analysed(m)],
        "modalities_skipped":   [m for m in MODALITIES
                                 if not subject.analysed(m)],
        "languages_detected":   subject.languages,
        "languages_unsupported": [l for l in subject.languages
                                  if l not in SUPPORTED_LANGUAGES],
    }
# The rule that makes this useful:
# NEVER return "no findings" on its own. Return
#   {"findings": [], "completeness": {...}}
# so a reviewer can tell a clean subject from an invisible one.
#
# And note what is NOT in this output: any protected characteristic.
# Counts and coverage only. A completeness score must not become
# a side channel for the data the report is supposed to redact.

That last comment matters. A completeness field such as "3 items withheld as non-job-relevant" is fine. "3 items relating to religion" is a disclosure and defeats the redaction the report exists to perform, as covered in what social screening is.

What depth should you expect?

Depth has three dimensions and vendors usually quote one. Ask about all three, and ask for the figure achieved rather than the figure offered.

DimensionTypical offerThe question to ask
Historical lookback7 to 10 years of post historyWhat lookback was actually achieved on this subject? Platform limits often bite before your window does
Content volumeUnstatedHow many items were reviewed? A screen of 40 posts and a screen of 4,000 both produce a report
Signal breadthA list of behavioural categoriesWhich categories were applied, and were they the ones in my policy? Consistency is only provable if the policy version is recorded

On the first row, note the asymmetry with a background check. Court records are permanent and retrievable. Social content is neither, so a 10-year lookback offer is a ceiling rather than a promise, and the achieved figure is the one that belongs in the file.

Why does coverage now have a regulatory dimension?

Because in immigration it stopped being a product choice. From 30 March 2026 the United States expanded mandatory social media screening to more than fifteen visa categories, covering both nonimmigrant and immigrant applications including H-1B, H-4, L-1, O-1, E-1 and E-2, TN, F-1, J-1, the EB-1 through EB-3 categories, K-1 and family-based green cards.

Three coverage consequences follow directly.

  • Declared handles create a checkable denominator. DS-160 and DS-260 require disclosure of handles used over a five-year lookback, including deleted accounts. A screen that cannot locate a declared handle has a measurable gap, not an unknown one.
  • Regional platform coverage becomes mandatory rather than nice to have. The free-text field exists because applicant populations use platforms that are not in any global top ten.
  • Pseudonymous account discovery is now consequential. The July 2025 refusal over an undisclosed public Reddit account shows what happens when a findable account is not surfaced.

If you build in this space, the requirement has moved from "screen social media" to "screen the specific accounts a subject declared, across the platforms they actually used, in the languages they actually posted in." That is a coverage specification, and it is far more demanding than a platform count.

What should you ask a vendor?

  1. Give me the platform list, filtered to my subject population's countries. Not the global count. A number without a list is marketing.
  2. How do you find accounts a subject did not declare? And what is your false positive rate on common names.
  3. Does your report distinguish clean from invisible? Show me a sample report on a subject with a private account.
  4. Which modalities do you analyse, and at which tier? Specifically: is audio and video included in my plan, or is it a premium feature.
  5. Which languages are natively classified, versus translated first? And what is measured performance per language.
  6. What lookback did you actually achieve on the last hundred subjects, against what you offer.
  7. How many content items were reviewed on a typical subject in my population.
  8. What happens on an ambiguous identity match, and does a human review it before it reaches the report.

Then run the test the answers invite: screen a subject you know, ideally a colleague who consents, including one with a private account and one who posts in a second language. An hour of that reorders most shortlists, and it is the same principle as the field-level test in our social data API coverage post.

Where does Phyllo fit, and where does it not?

We are the data layer beneath screening products rather than a report vendor. Phyllo's social screening supplies structured signals across social platforms with identity resolution and account linkage underneath, so the product you build spends its engineering on classification, redaction and completeness reporting rather than on collection and account discovery. It sits under background verification for BGV firms, influencer vetting for brand safety, and visa and immigration checks. Coverage is published per platform at getphyllo.com/coverage, and we hold GDPR compliance and SOC 2 Type II.

Where we are the wrong purchase, plainly. If you need a finished FCRA-compliant report with redaction applied, an adverse action workflow attached and analyst review on ambiguous results, buy from a consumer reporting agency. Ferretly, Checkr, Sterling and others do exactly that and we are not a substitute. We are the right purchase when screening is what your product does at volume and you need the layer underneath. Costs for both models are in social screening pricing.

What does coverage mean for a social screening API?

Three things reported as one: reach, whether the accounts can be found; visibility, whether the content is public; and comprehension, whether the tool can classify it.

Can social screening see private accounts?

No, and there is no legitimate route around a privacy setting. A private account should be reported as found but inaccessible, a coverage gap rather than a clean result.

What does no findings mean on a screening report?

Either the subject is clean or nothing was visible, and most reports do not distinguish. Ask for accounts found, accounts accessible, items reviewed and lookback achieved.

Does social screening analyse video?

Often not, or only at a premium tier. Short-form video dominates TikTok, Reels and Shorts, so a tool reading only captions reports on descriptions rather than content.

Does social screening work in languages other than English?

Performance varies, and support is claimed more broadly than it is measured. Slurs are region specific, so an English-trained classifier misses more content elsewhere.

How many platforms should a screening tool cover?

Enough for where your subjects actually post, rarely the global top ten. Ask for the list filtered to their countries. The DS-160 anticipates Sharechat, Koo and Moj.

How far back does social screening look?

Seven to ten years is commonly offered, but that is a ceiling, because platform limits often bite first. Ask for the lookback actually achieved on recent subjects.

Is social media screening required for US visas?

From 30 March 2026 the US expanded mandatory screening to more than fifteen visa categories. Applicants disclose handles over a five-year lookback, including deleted ones.

Table of Content
See Phyllo in action
  • No Credit card required
  • GDPR & SOC2 Type II
  • 30-min Onboarding
Book a Demo

Be the first to get insights and updates from Phyllo. Subscribe to our blog.

Ready to get started?

Sign up to get API keys or request us for a demo