A field-by-field split of what public and authenticated social data can and cannot return, why estimates are not measurements, and which survives an audit.
- Public data is what a platform shows a stranger, authenticated social data is what it shows the account holder, and that line fixes your fields.
- Public can never return 5 things: true impressions, audience demographics, saves and shares, expiring-format performance, and earnings.
- Authenticated cannot return 6 things: unconnected creators, most pre-connection history, competitors, post-revocation data, day 1 breadth, and API-withheld fields.
- Public estimates are models, not measurements; an API that never returns null is either always confident or never checking.
- Run both with precedence: authenticated wins on any field it returns; never average or sort by recency across sources.
Public social data is anything a platform renders to someone who is not logged in and has no relationship with the account: handle, bio, follower count, captions, public like and comment counts. Authenticated data is what the platform renders to the account holder once they have granted your application permission: true impressions, audience demographics, performance on expiring formats, earnings.
That is the whole distinction, and it is not a difference of degree. No amount of money, coverage or engineering moves a field from one side to the other, because the boundary is a platform policy decision rather than a technical limitation.
What follows is the full picture in both directions, field by field and platform by platform. What public returns and never will. What authenticated returns and, in the section most articles skip, what it cannot. How the estimates are actually built. What happens on revocation, which source survives an audit, and how to run both without corrupting your data.
What does authenticated actually mean?
It means the account holder completed an OAuth flow in your product and approved a specific set of scopes, and the platform now answers your requests on their behalf. 4 things exist afterwards that did not before, and each has consequences later.
- A consent record. A timestamped, auditable fact that this person authorised this application for these scopes. This is the artefact that matters in a dispute.
- A scoped token. Permission is granular. Approving profile read does not grant insights, and requesting scopes you do not use is a common app review rejection.
- A revocation path. The creator can withdraw permission from their platform settings at any time, without telling you.
- A lifecycle you now own. Tokens expire. Instagram long-lived tokens last about 60 days, LinkedIn access tokens 60 days, TikTok access tokens 24 hours. None of them refresh themselves.
Public data has none of the 4, because there is no relationship. That absence is the source of both its advantages and its liabilities.
What can public social data return?
More than it gets credit for, and it is the right tool for a large class of products. Anything visible in a logged-out browser is available at scale, on accounts that have never heard of you.
| Field group | Typical contents | Good for |
|---|---|---|
| Profile | Handle, name, bio, avatar, external link, follower and following counts, post count | Discovery, screening, benchmarking |
| Content | Captions, hashtags, media URLs, permalinks, timestamps, media type | Listening, campaign tracking, content analysis |
| Public engagement | Like counts, comment counts, video and Reels view counts | Ranking, filtering, trend detection |
| Derived | Engagement rate, posting frequency, growth rate, caption length | Shortlisting at volume |
| Contact | Public emails and links where the creator published them | Outreach and enrichment |
The strength is breadth and immediacy. You can query millions of accounts on day 1 with no relationship to any of them, which is not a compromise but a different capability. We covered the official public surfaces and their limits in social media public data.
What can public data never return?
5 field groups. This is a structural boundary rather than a gap that will close, because each one is a deliberate decision to release the data to the account holder alone.
| Field group | Why it is withheld | What public tools offer instead |
|---|---|---|
| True impressions and reach | Rendered only inside the creator's own analytics | Estimated reach, modelled from followers and engagement |
| Audience demographics | Age, gender and location released to the owner only | Inferred demographics from follower sampling |
| Saves, shares, profile visits | Private engagement signals never exposed publicly | Likes and comments as a proxy, which misses the formats driving discovery |
| Expiring content performance | Stories and similar formats vanish in 24 hours and are never publicly archived | Nothing. There is no public layer for them |
| Earnings and payouts | Never rendered on any public surface on any platform | Rate card estimates, which are asking prices rather than records of payment |
The right-hand column is the part worth arguing about, so the next section takes it apart.
To check which side of the line a specific field falls on for your platforms, the per-platform list is public. See the coverage list
How are public estimates actually derived?
By modelling from the public signals available. That is legitimate and useful, and it becomes a problem only when a modelled number is presented, resold or filed as a measured one.
The 2 most common derivations work roughly like this. Estimated reach extrapolates from follower count and recent engagement, often with a platform-specific multiplier. Inferred audience demographics sample a subset of visible followers, classify each from their own public profile, and project that distribution onto the whole audience.
Both can be reasonably accurate at the top of the follower distribution and both degrade sharply on small accounts, where the sample is thin and the multiplier is unstable. If your product serves nano and micro creators, that is exactly where you need the numbers to hold and exactly where they do not.
One question sorts careful vendors from careless ones: what does the API return when the model cannot produce a defensible number? Null is the correct answer. A vendor who always returns a number is either always confident or never checking, and only one of those is possible. Ask 2 follow-ups in the same call: is this field measured or modelled, in writing, and is the model documented. A proprietary model is fine. An undisclosed one you are reselling as measurement is not.
What can authenticated data return that public cannot?
The same 5 groups, as measurements rather than models, because they come from the platform's own analytics through its official channel.
| Field group | What arrives | Why it matters commercially |
|---|---|---|
| True impressions and reach | Exact counts per post from the platform | Real cost per impression instead of a model. This is the number a media plan is defensible on |
| Audience demographics | Age, gender and location distributions | A brand can rely on it. Inferred demographics do not survive a client challenge |
| Expiring formats | Performance captured while the content is live | A large share of campaign delivery happens in formats public data cannot see at all |
| Earnings and payouts | Where the platform exposes them | The only route to income verification. No public source holds this at any size |
| Verified cross-platform identity | One creator, their connected accounts, confirmed | Public identity linking is inference. Consent makes it a fact |
The mechanism matters as much as the fields. Because the data travels from the platform to you through an authorised channel, you get 3 properties public collection cannot offer: it is first-party rather than modelled, auditable back to a consent record, and carries a documented legal basis.
Which scope unlocks what, platform by platform?
The boundary is not abstract. Each platform names the permission that crosses it, and knowing which one saves a scoping conversation.
| Platform | Demographics | Earnings | Notes |
|---|---|---|---|
instagram_manage_insights | Not exposed | The only scope returning audience breakdowns, and only for the authorising account | |
| TikTok | Not available | Not exposed | The native developer API has no audience demographics scope at all |
| YouTube | Via analytics scopes | yt-analytics-monetary.readonly | The monetary scope is the only revenue path and is a superset of the activity scope |
| Withheld at every tier | Not exposed | Creator-level demographics are not available on any commercial product |
Scope-by-scope references for each: Instagram, TikTok, YouTube and LinkedIn.
What can authenticated data not return?
6 things, and this is the section that usually goes missing. Better to find them here than in month 3.
- Anyone who has not connected. The whole constraint. A creator who never authorised your app does not exist in your authenticated data at any price. Discovery, prospecting and competitor research cannot be served this way.
- Most history from before they connected. Platforms expose a limited analytics window, so connecting today does not hand you 3 years of impressions. Some fields backfill, many do not. Ask per field, per platform.
- Anything after revocation. The moment a creator disconnects, the tap closes.
- Breadth. Your authenticated coverage equals your connected user count. On day 1 that is zero.
- Fields the platform withholds from the owner through its API. The honest ceiling on the model. Consent gets you what the account holder can see through the API, not what they can see in the app. Historical Stories are the clean case: after 24 hours the creator cannot retrieve them either.
- Immunity from platform limits. An authenticated layer inherits every rate limit, quota and deprecation. TikTok tokens expire every 24 hours. Google revokes YouTube-scoped refresh tokens when a user changes their account password. Meta retired an entire API with 90 days notice.
Point 5 is worth sitting with when evaluating any consented vendor, including us. The right question is not "do you have consent" but "which fields does the platform expose to an authorised app, per platform." Those are different lists, and only the second one is coverage.
What happens when a creator revokes access?
Authenticated data stops immediately, usually with no warning, and you may owe deletion depending on your lawful basis and what you told the creator. Public data has no revocation mechanism, which sounds like an advantage and cuts both ways.
| Authenticated | Public | |
|---|---|---|
| How it ends | Creator revokes in platform settings, or the token expires | It does not. There is no relationship to sever |
| How you find out | Calls start failing. You will not get a notification | Not applicable |
| Historical data you hold | Retention depends on your basis and your notice | Unaffected by any creator action |
| Practical consequence | A live connection becomes a dead one and your interface must handle it | You keep collecting, and you keep owning the compliance question |
| Detectability | High. A failed call is loud | Low. Nothing tells you a record is stale |
Build for revocation explicitly. Distinguish a revoked connection from an expired token in your error handling, because one needs the creator to reconnect and the other needs your refresh job to run. Surface disconnected accounts rather than showing stale numbers. The full architecture is in consent and data retention for authenticated social data.
Which source survives an audit?
Authenticated, and it is not close. This decides procurement in regulated categories, and it has almost nothing to do with data quality.
The question an auditor, a client or a regulator asks is not "is this number accurate" but "where did it come from, and what entitled you to have it." Authenticated access answers both: the platform returned it, and here is the timestamped consent record that authorised the request. Public collection answers neither with the same force.
| Question | Authenticated | Public |
|---|---|---|
| Where did this figure come from? | The platform, through an authorised channel | A vendor, by collection or by model |
| Can you evidence the basis? | Yes, a consent record with scopes and a timestamp | A general argument about public availability |
| Is it measured or modelled? | Measured | Often modelled, and not always labelled |
| Can the subject object? | They can revoke, which is itself the answer | There is usually no mechanism they know of |
| Fit for hiring, lending, visas | Designed for it | Depends heavily on jurisdiction and use |
If you sell into fintech, HR, background verification or immigration, this table is why a consent-first model is a procurement requirement rather than a preference. It is not legal advice, and the position varies by jurisdiction and use case.
Why do the 2 have opposite coverage curves?
Because one starts complete and erodes, and the other starts empty and compounds. This is the most useful thing here for anyone making a 3-year architecture decision.
Public coverage starts at its maximum on day 1 and decays. You can query everything visible today. Then platforms restrict. Meta retired the Instagram Basic Display API entirely. Instagram hashtag search will not return a username. TikTok's Research API narrowed to academic institutions. LinkedIn withholds creator engagement at every commercial tier. Every one of those moves went the same direction and none reversed.
Authenticated coverage starts at zero and grows with your product. Every creator who connects adds permanent depth on that creator, and nothing a platform does about public access takes it away, because the account holder authorised it. Your coverage curve is your adoption curve.
| Public | Authenticated | |
|---|---|---|
| Day 1 | Maximum breadth, zero depth | Zero breadth, maximum depth per connection |
| Direction of travel | Narrowing. Every platform change in 5 years has restricted access | Widening with your user base |
| What a policy change does | Can remove a field or an entire API overnight | Can change the API surface, not your permission |
| Risk concentration | Platform policy risk | Adoption and onboarding friction risk |
Neither curve is better. They mean different things for a business plan. If your roadmap depends on public breadth staying where it is today, that is a bet against a 5-year trend. If it depends on authenticated depth, the bet is on your own onboarding conversion, which you control.
How do you run both without corrupting your data?
With an explicit precedence rule. Most production systems use public data for discovery and screening, then authenticated data for verified numbers once a creator is onboarded. The failure is not in either source. It is in the merge.
- Authenticated wins on any field it returns. Not the newer value, not the higher value. The source.
- Never average or reconcile numerically. A modelled reach of 40,000 and a measured reach of 26,000 do not average to anything meaningful.
- Carry provenance on every field, all the way to the interface. If a brand asks whether a number is measured or estimated, the answer must be in your data model.
# Field-level resolver with provenance. Precedence, not arithmetic.
SOURCE_RANK = {"authenticated": 2, "public": 1} # higher wins
def resolve(field, candidates):
usable = [c for c in candidates if c["value"] is not None]
if not usable:
return {"value": None, "source": None, "measured": None}
best = max(usable, key=lambda c: (SOURCE_RANK[c["source"]],
c["fetched_at"]))
return {
"value": best["value"],
"source": best["source"],
"measured": best["measured"], # render this in the UI
"fetched_at": best["fetched_at"],
"alternates": [c for c in usable if c is not best],
}
resolve("reach", [
{"value": 40000, "source": "public", "measured": False,
"fetched_at": "2026-07-27T09:00Z"},
{"value": 26410, "source": "authenticated", "measured": True,
"fetched_at": "2026-07-27T08:00Z"},
])
# -> 26410, authenticated, measured=True
#
# The authenticated value is OLDER and still wins. Recency is a
# tiebreaker WITHIN a source, never across sources. Sorting by
# freshness globally silently promotes an estimate over a
# measurement, and nothing in your logs records that it happened.
That last comment is the rule people get wrong most often, and it is invisible when it fails. On identity, reconciling a public record to a consented one is its own problem, and getting it wrong means attributing one creator's numbers to another. That is what identity resolution exists for.
Which should you build on?
Answer one question: does your product require the creator to show up? If they are signing in anyway, authenticated access adds no friction and unlocks fields no index holds. If your product must work on people who have never heard of you, authenticated access cannot help at any price.
| What you are building | Primary source | Why |
|---|---|---|
| Creator discovery or search | Public | Needs breadth on people who have not signed up |
| Media kit or creator dashboard | Authenticated | The creator is logged in and wants their real numbers |
| Income verification or creator lending | Authenticated | No public source holds earnings at any size |
| Brand monitoring and social listening | Public | There is no account holder to ask |
| Influencer vetting before spend | Both, in sequence | Public to screen, authenticated to verify once onboarded |
| Social screening for hiring or visas | Public, with a consent process | The subject rarely connects. Consent is a legal design question |
| Campaign reporting to a client | Authenticated | Estimated reach does not survive a client challenge |
| Competitor benchmarking | Public | Nobody authorises you to study them |
Where does Phyllo fit?
We are the authenticated side. A creator connects through your product, and Phyllo's social data API returns true impressions, audience demographics, engagement across every format including expiring ones, and earnings where the platform exposes them, normalised across 25+ platforms through one integration with the token lifecycle handled. Per-platform, per-field coverage is public at getphyllo.com/coverage and the API reference needs no sales call.
We also run social listening and social screening on public signals, because those jobs have no account holder to ask. Both models exist here for the reasons in this post rather than as a product accident.
Where we are the wrong purchase, stated plainly. Everything in the public column above is available on creators who have never heard of you, and we cannot serve that, because we need the creator to connect. If discovery, competitor research or cold breadth is the job, buy a public index. Modash for creator discovery with audience modelling, EnsembleData for TikTok volume, Bright Data or Oxylabs for enterprise scale. Those are recommendations rather than politeness, and the cost models for both are compared in social data API pricing.
The short version
Write down the fields your product renders, then mark each one public or authenticated using the tables above. That exercise takes 20 minutes and settles more architecture than any vendor evaluation, because it tells you which category of product you are shopping in before you talk to anyone.
If your fields are all in the public column, buy a public index and do not pay for consent. If any of them are in the authenticated column, no index will ever return them, and that is worth establishing in week 1 rather than month 6.
Check which of the fields you need require creator authorisation, or talk to us about your connect flow. Get a demo
What is the difference between authenticated and public social data?
Public data is what a platform shows someone not logged in: handle, bio, follower count, captions, public likes. Authenticated data is what it shows the account holder after they authorise your app.
Can I get audience demographics without creator consent?
Not as measurements. Every major platform releases audience age, gender and location to the owner only. Public tools infer demographics by sampling visible followers, which is directional at best.
What can authenticated social data not give me?
Anyone who has not connected, most history from before they connected, competitor data, anything after revocation, breadth on day 1, and fields the platform withholds from the owner through its API.
Is authenticated data more accurate than public data?
For fields both return, authenticated is measured rather than modelled, which makes it more defensible. For creators who never connected, authenticated data is not less accurate, it is absent.
Can I use both public and authenticated data together?
Yes, most production systems do: public for discovery and screening, authenticated for verified numbers after onboarding. Authenticated wins any field it returns; never average across sources.
Which is better for compliance?
Authenticated, because the consent record documents where a figure came from and what entitled you to request it. Public collection is more contested and varies by jurisdiction. Not legal advice.
Does consent unlock every field the creator can see?
No. Consent unlocks what the platform exposes to an authorised app, which is narrower than what the creator sees in the app. Instagram Stories older than 24 hours are out of reach for the creator too.
What happens to my data when a creator revokes access?
New reads stop at once, usually with no notice, so you find out when calls fail. Keeping historical data depends on your lawful basis. Distinguish a revoked connection from an expired token.




