Why signal-to-outcome correlation is not computable for social screening, and a taxonomy built on attribution, recency, ambiguity and job-relatedness.
- Nobody can publish signal-to-outcome correlation for social screening: flagged candidates are not hired, so their outcomes are never observed, and more volume cannot fix that.
- Rank signals by evidentiary quality instead, on 4 dimensions: attribution, recency, ambiguity and job-relatedness.
- Associational signals are the biggest trap: a follow is not conduct, and it tracks protected characteristics more reliably than behaviour.
- Volume is not validation: a vendor citing millions of screens has told you about their sales, not their accuracy.
- Not legal advice: validation requirements and protected categories vary by jurisdiction and use case.
The honest answer to the question in that title is that nobody knows, and the reason is worth understanding before you buy or build anything in this category.
To measure whether a screening signal predicts risk, you would need to observe outcomes for the people it flagged. But the point of screening is to prevent those people being hired, so the outcome never happens. You observe outcomes only for candidates the screen let through, which is a filtered population selected on the very thing you are trying to measure.
This is a structural problem rather than a data-collection problem, and no amount of screening volume fixes it. So a defensible taxonomy has to be built on something you can actually assess: the evidentiary quality of each signal. What follows is that taxonomy, the 4 dimensions that determine it, and the category of signal that creates legal exposure regardless of how well it appears to work.
Why can nobody measure this?
Because screening destroys the counterfactual. Work through what you would need.
| Group | Outcome observable? | Why |
|---|---|---|
| Flagged, not hired | No | They never worked for you. Nothing to observe |
| Flagged, hired anyway | Rarely | A small, unusual group. Someone overrode the flag for a reason |
| Not flagged, hired | Yes | The only group with clean outcome data |
| Not flagged, not hired | No | Declined for unrelated reasons |
Only the third row produces usable data, and it is defined by having passed the screen. Correlating signals against outcomes within it tells you how signals behave among people the signals cleared, which is not the question.
Lending has the same shape and a name for it: reject inference. The partial fix there is to approve a random sample of rejected applicants and observe what happens, which is expensive but possible because the cost of a bad loan is bounded and quantifiable. No employer will hire a random sample of candidates their screening flagged for threats or harassment, and no one should ask them to. So the experiment that would validate these signals is one nobody will ever run.
Which leads to a practical rule for vendor conversations. If a vendor presents correlation data between social signals and employment outcomes, ask which population it was computed on. If the answer is hired employees, it was computed on survivors. If there is no answer, the number was not computed at all.
What should the taxonomy be built on instead?
Evidentiary quality, which is assessable from the signal itself, and which is also what a dispute will turn on. 4 dimensions.
| Dimension | The question | Why it decides quality |
|---|---|---|
| Attribution | Is this definitely the subject? | The whole report collapses if it is the wrong person. Carry match confidence into the finding |
| Recency | When did it happen? | A statement from 9 years ago and one from last month are different facts about a person |
| Ambiguity | Does it need interpretation? | An explicit threat means 1 thing. A sarcastic reply means whatever the reader decides |
| Job-relatedness | Does it bear on the role? | The only basis on which a finding can lawfully inform a decision |
A signal scoring well on all 4 is reportable. A signal scoring badly on any 1 of them is a liability, and the common failure is a signal that scores well on 3 and badly on attribution, because that one looks like evidence right up until it turns out to be somebody else.
What does the taxonomy look like?
| Tier | Signal type | Evidentiary quality | Handling |
|---|---|---|---|
| 1 | Explicit statements: direct threats, admissions of illegal conduct, targeted harassment naming a person | Strong. Unambiguous, attributable, conduct-based | Report with excerpt, date and link |
| 2 | Contextual conduct: aggressive exchanges, disparagement of an employer, content that needs surrounding context to read | Moderate. Attributable but interpretation-dependent | Report with context. Human review required |
| 3 | Associational: follows, likes, group membership, shares without comment | Weak. Not conduct. Frequently proxies protected characteristics | Do not report as a finding |
| 4 | Protected characteristics: race, religion, age, disability, pregnancy, orientation and the rest | None. Unlawful to consider | Drop before any human sees it |
The tier boundaries are doing the work. Tier 1 and Tier 2 are both conduct, and they differ in how much reading the reviewer has to do. Tier 3 is not conduct at all, and that is the distinction most screening products blur.
Tier 4 handling is covered properly in what social screening is and the pipeline position in how to build a social screening product. The short version: drop, never flag, and report only a count, because "3 items withheld relating to religion" discloses the thing the redaction existed to hide.
Why are associational signals the biggest trap?
Because they are the cheapest to collect, the easiest to count, and the weakest thing in the report.
A follow is not an action taken against anyone. People follow accounts to monitor them, to argue with them, because a friend did, or because an algorithm suggested it and they tapped without much thought. Treating a follow as conduct requires an inference that the subject cannot contest and the reviewer cannot verify.
The more serious problem is correlation with protected characteristics. Consider what a follow graph actually encodes.
| Associational signal | What it frequently proxies |
|---|---|
| Follows accounts in a particular language | National origin |
| Follows religious organisations or clergy | Religion |
| Follows condition-specific support communities | Disability or medical condition |
| Follows pregnancy or fertility accounts | Pregnancy |
| Follows advocacy organisations | Political affiliation, and sometimes orientation or race |
| Follows accounts skewing to an age cohort | Age |
A neutral-looking signal that reliably predicts a protected characteristic is a proxy for it, and acting on the proxy is indirect discrimination. That holds whether or not anyone intended it, and whether or not the model was told what the characteristic was. A screening product that scores candidates on their follow graph has built a protected-characteristic classifier by accident and will struggle to prove otherwise.
This is why Tier 3 belongs out of the report rather than low in it. A weak signal that is also a discrimination vector is not worth the small amount of information it carries.
Why does signal quality fall as collection gets easier?
Because the signals that are cheap to gather are the structured, countable ones, and the ones that carry evidentiary weight are unstructured and expensive to find.
| Signal | Collection cost | Evidentiary weight | Tempting because |
|---|---|---|---|
| Follow and like counts | Very low | Very low | Countable, chartable, easy to score |
| Keyword matches in captions | Low | Low to moderate | Simple, fast, language-dependent |
| Classified text content | Moderate | Moderate to strong | Mature classifiers |
| On-screen text in images | Higher | Moderate to strong | Needs OCR, often skipped |
| Spoken content in video | Highest | Strong | Needs transcription. Frequently gated to premium tiers |
That inversion creates a specific commercial pressure: a product can look comprehensive by counting cheap signals while missing the expensive ones that actually matter. Short-form video is the median post on the platforms most subjects use, so a caption-only screen is reading the description rather than the content. Coverage mechanics are in social screening API coverage.
The test for any screening product: ask what proportion of reviewed items were video, and whether the audio was analysed. A vendor that cannot answer is scoring the cheap signals.
How should a finding be scored?
By carrying the 4 dimensions into the record rather than collapsing them into a severity number. A single score hides exactly the information a dispute needs.
# A finding is not a score. It is 4 assessments plus the evidence.
finding = {
"tier": 1, # explicit conduct
"category": "threat_of_violence",
"excerpt": "...", # the evidence itself
"permalink": "...",
"posted_at": "2026-03-14",
# The 4 dimensions, carried forward, never collapsed
"attribution_confidence": 0.94, # from identity resolution
"age_days": 135,
"requires_interpretation": False,
"job_related": True, # mapped to a policy category
"policy_version": "screen-policy-v7",
"reviewed_by": "human",
}
# Rules that follow from keeping them separate:
# - attribution_confidence below threshold -> human review, never auto-flag
# - requires_interpretation True -> human review, mandatory
# - job_related False -> not a finding. Drop it.
# - tier 3 or 4 -> never reaches this structure
#
# A single 0-100 "risk score" destroys all of this. It cannot be
# disputed, cannot be explained to a candidate, and cannot be
# defended when somebody asks what it was made of.
The last comment is the design argument. A composite score feels like a product and behaves like a liability, because the moment a candidate or a regulator asks what produced it, you need the components back, and a model that blended them cannot return them.
What about recency?
It deserves an explicit policy rather than an implicit one, because an unbounded lookback produces findings that nobody intended to act on.
Content from a subject's teenage years, on an account they have had since school, is a different fact from content posted last quarter. Several jurisdictions limit how far back consumer reports may reach for certain information, and even where no limit applies, a 10 year old post is weak evidence about a person's current conduct.
3 things to decide and write down.
- Lookback window per category. Threats and illegal conduct may warrant a longer window than disparagement.
- Whether age reduces tier. A Tier 1 finding from 8 years ago may be reportable with its age stated prominently rather than suppressed.
- Achieved versus intended lookback. Platforms limit history, so record how far back you actually reached rather than how far your policy allows.
That third point matters for report integrity. A policy of 7 years and an achieved reach of 14 months is a gap the reader should see.
What does a vendor need to answer?
- Which signal tiers do you report, and do you report associational signals as findings? If yes, ask how they handle the protected-characteristic correlation.
- Do you produce a composite risk score? If so, can you decompose it on request. If not, it cannot be disputed.
- What is your attribution confidence distribution, and what happens below threshold? Human review is the only acceptable answer.
- What proportion of reviewed items were video, and was audio analysed? This separates comprehensive products from cheap ones.
- How do you validate your categories? Honest answer: against reviewer agreement and policy mapping, not against employment outcomes. A vendor claiming outcome validation should be asked which population.
- Which languages have native classifiers rather than translation? Applying an English-tuned model to other languages applies a different standard to different candidate populations.
What can be measured?
Plenty, just not predictive accuracy. These are the metrics a screening product can honestly report, and they are what a serious buyer should ask for.
| Measurable | What it tells you |
|---|---|
| Inter-reviewer agreement | Whether 2 humans reading the same item reach the same categorisation. Low agreement means the category is badly defined |
| Attribution accuracy | Verifiable against subject-confirmed accounts. The one place a hard accuracy number is achievable |
| Coverage completeness | Accounts declared versus found versus accessible, items reviewed, lookback achieved, modalities analysed |
| Classifier precision against reviewed items | Of items flagged, how many a human upheld. Measurable without any outcome data |
| Language and modality parity | Whether flag rates differ across languages in ways the content does not explain |
| Appeal and correction rate | How often a finding is overturned on dispute |
Row 5 is the one to instrument early and the one most likely to surface a fairness problem before someone else finds it. If your flag rate on Portuguese content is materially different from English and the content is not, the difference is in the classifier rather than the candidates.
Where does Phyllo fit?
We supply the layer underneath the taxonomy rather than the taxonomy itself. Social screening provides the structured signals and identity resolution provides the attribution confidence that Tier 1 and Tier 2 findings depend on, so your product spends its engineering on categorisation, redaction and reporting rather than on account discovery and collection.
Attribution is the part of this worth buying rather than building, because it is the dimension that invalidates everything downstream when it is wrong, and its accuracy comes from cross-platform data rather than from better code. It sits under background verification for BGV firms and influencer vetting for brand safety. Per-platform coverage is at getphyllo.com/coverage and the API reference is public.
What we will not sell you is a validated risk score. Nobody can validate one against employment outcomes for the reason at the top of this page, and a vendor offering one has either not thought about the problem or is describing something else. Cost models for the category are in social screening pricing.
The short version
Stop asking which signals predict risk, because the question cannot be answered and a vendor who answers it confidently has told you something about their methodology. Ask instead which signals would survive a dispute: attributable to this person, recent enough to matter, unambiguous without interpretation, and related to the job.
Then keep associational signals out of the report entirely. They are the cheapest thing to collect, the weakest thing in evidence, and the most reliable proxy for characteristics you are forbidden to consider. That combination is not a trade-off worth making.
Want the attribution layer that Tier 1 and Tier 2 findings depend on? Get a demo
Which social media signals predict employment risk?
Nobody can measure which signals predict risk: flagged candidates are not hired, so the only clean outcome data belongs to people the screen cleared. Rank signals by evidentiary quality instead.
Why can screening vendors not publish accuracy data?
The validation experiment cannot be run. Lending's partial fix approves a random sample of rejected applicants; no employer will hire a random sample of candidates flagged for threats or harassment.
Should a screening report include who a candidate follows?
No. A follow is not conduct, and follow graphs strongly track protected characteristics such as religion, national origin, disability and pregnancy. Acting on that proxy is indirect discrimination.
Is a composite risk score useful?
It is convenient and hard to defend. A single number cannot be disputed, explained to a candidate or decomposed for a regulator. Keep attribution, recency, ambiguity and job-relatedness separate.
What can a screening product actually measure?
Inter-reviewer agreement, attribution accuracy, coverage completeness, classifier precision against human-upheld flags, language and modality parity, and appeal rates. None needs outcome data.
How far back should a screen look?
Set a lookback per category and write it down, then record how far back you actually reached, since platforms limit history. A 7-year policy with 14 months achieved is a gap the reader should see.
Does a vendor's screening volume indicate accuracy?
No. Volume indicates sales. Accuracy on the only dimension where a hard number is achievable, attribution, is measurable against subject-confirmed accounts, and that is the figure worth asking for.



