Refresh economics force every large public social data index to tier, so the median record is months old. The arithmetic, and how to measure your own.
- Uniform daily refresh of a 350 million profile index costs about $127,750,000 a year, so every large index tiers.
- Tiering puts the median record at about 160 days old and the p95 at about 344 days, modelled rather than measured.
- Unlike public social data, authenticated data is fetched live at call time, so its only staleness is your cache.
- Fields decay at different rates: a bio holds for months, but most post engagement lands within 24 to 48 hours.
- Measure any vendor in an afternoon: make a controlled change, poll until it appears, and report p50, p95 and never-propagated counts.
When a discovery tool shows you a creator with 240,000 followers, the useful question is not whether the number is right. It is when it was true.
That question has a structural answer before it has an empirical one. Refreshing a large public index is expensive enough that no vendor can keep every record current, so they refresh by tier: a small set frequently, the long tail rarely. That is a sound engineering decision and it has an unavoidable consequence for the age of a record you pull at random.
Below: the arithmetic that forces tiering, the age distribution it implies, why authenticated data sits outside the problem entirely, and a method you can run against any vendor to get measured figures rather than modelled ones.
Why can no large index be uniformly fresh?
Because the cost scales with records multiplied by frequency, and both numbers are large. Take an index of 350 million profiles, which is the scale the larger discovery products advertise, at a collection cost of $1 per 1,000 records, which is the low end of published rates.
| Refresh cadence | Annual cost | Verdict |
|---|---|---|
| Daily | About $127,750,000 | Larger than most companies' entire revenue |
| Weekly | About $18,200,000 | Not a business |
| Monthly | About $4,200,000 | Exceeds most infrastructure budgets outright |
| Quarterly | About $1,400,000 | Affordable, and the data is a quarter old |
So uniform refresh at any cadence a buyer would call fresh is unaffordable. Every vendor operating at this scale tiers by demand: profiles that appear in results often get refreshed often, and the long tail does not. We worked through that design in how to build an influencer discovery product.
This matters because it changes the nature of the question. Freshness is not a quality difference between vendors who care and vendors who do not. It is a constraint that applies to all of them, and the only variable is how they distribute the staleness.
What age does tiering imply?
Take a representative tier structure and the age distribution falls out of it directly. A record's age at any moment is roughly uniform between zero and its tier's cadence.
| Tier | Share of index | Cadence | Mean age of a record |
|---|---|---|---|
| Hot | 0.1% | Daily | 0.5 days |
| Warm | 1% | Weekly | 3.5 days |
| Cool | 10% | Monthly | 15 days |
| Cold, the long tail | 88.9% | Annual | 182.5 days |
Because the cold tier holds 88.9% of the index, a record pulled at random is overwhelmingly likely to come from it. Working the percentiles through:
| Percentile | Age | What this means |
|---|---|---|
| p50 | About 160 days | The median record is over 5 months old |
| p75 | About 262 days | A quarter of records are older than 8.7 months |
| p90 | About 324 days | A tenth are older than 10.8 months |
| p95 | About 344 days | 1 in 20 records is nearly a year out of date |
| Mean | About 164 days | 5.5 months |
These are modelled figures, not measured ones. They describe what a representative tier structure implies, and the tier structure itself is forced by the refresh economics above. The useful claim is not that any specific vendor sits at exactly these numbers. It is that a vendor operating a large public index at a sustainable cost cannot be far from this shape, because the alternative costs 123 times more.
Which gives you a question worth asking any vendor: what proportion of your index was refreshed in the last 30 days? A vendor who tiers well will answer with a number in the low tens of percent. A vendor who claims a high figure across a large index is describing a cost structure that does not add up.
If freshness matters for a specific field, check first whether it can be read live rather than indexed. See the coverage list
Why does authenticated data not have this problem?
Because there is no index. An authenticated read goes to the platform at the moment you call it and returns the current value, so the record has no age beyond your own caching decision.
That is an architectural difference rather than a performance one, and it is worth being precise about what it does and does not mean.
| Public index | Authenticated read | |
|---|---|---|
| Where the value comes from | A snapshot stored at collection time | The platform, at call time |
| Age of the value | Whatever the tier cadence allows | Zero, plus your cache time-to-live |
| Who bears the refresh cost | The vendor, spread across all customers | You, per call, only for accounts you read |
| Scales with | Index size multiplied by cadence | Number of connected accounts you actually query |
| Cheap when | You read many accounts once | You read few accounts often |
| Stale when | Always, by design, for the tail | Only if you cache aggressively |
The fourth row is why the 2 models have opposite cost curves. Re-reading the same connected account costs almost nothing, so authenticated data gets cheaper per read as you refresh more. Re-reading the same indexed profile costs full price every time, so public data gets more expensive the fresher you want it. The full cost comparison is in social data API pricing.
The honest limit: authenticated reads only exist for accounts that authorised you. This is not an argument that one model is better, it is that freshness is a solved problem on one side and a permanent constraint on the other. The boundary between them is set out in authenticated versus public social data.
Why is a single freshness number meaningless?
Because fields decay at rates that differ by orders of magnitude, so an index that is 30 days old is badly wrong about some fields and perfectly correct about others.
| Field | Rate of change | Useful lifetime |
|---|---|---|
| Bio and display name | Rarely, a few times a year | Months. A 6 month old bio is usually still right |
| Profile picture | Rarely | Months |
| Follower count | Continuously | Hours to days, depending on account size |
| Post count | On every post | As long as the creator's posting interval |
| Engagement on a specific post | Fast, then asymptotic | Most engagement lands in the first 24 to 48 hours |
| Audience demographics | Slowly | Weeks to months |
| Ephemeral content metrics | Expires entirely | 24 hours, then the data does not exist |
The correct model is a time-to-live per field rather than per record, and it is worth implementing on your side even if your vendor does not expose one. A follower count from last quarter should be labelled differently from a bio from last quarter, because 1 of them is probably wrong and the other probably is not.
The engagement row carries a practical implication for anyone benchmarking creator performance. Because most engagement accrues in the first day or 2, a post collected a week after publication has close to its final numbers, while a post collected 6 hours after publication does not. Comparing the 2 without normalising for post age produces a ranking driven by collection timing rather than by performance.
What do the vendors say about their own freshness?
More than you might expect, and the honest ones say it in their documentation rather than their marketing.
- One connector's documentation states that none of its datasets use a date range and every run pulls all currently available data. That is a snapshot model stated plainly: there is no history to query, only a present to capture.
- Another tells users directly that its public data provides current profile snapshots and recent media, and that for deep historical analysis you should start tracking a profile early and let the tool build history over time.
- Several expose a fetch timestamp as a first-class column, such as a
data_fetched_atfield or an extraction timestamp. A vendor who ships that field has understood the problem. A vendor who does not is asking you to assume the data is current.
Read together, those 3 make the same point from different directions: public data is not a historical record, it is a recorder you start. You cannot backfill a competitor's growth curve from a standing start, and any vendor offering one has been collecting since then and is selling you their archive. We covered that in social media public data.
What has a hard expiry rather than an age?
Some data does not go stale, it disappears, and no refresh strategy recovers it.
| Data | Why there is nothing to refresh |
|---|---|
| Ephemeral content performance | Stories and similar formats expire after 24 hours and are never publicly archived. Capturing them means polling while they are live |
| Instagram hashtag recent media | The endpoint returns only media published within 24 hours of the query, so there is no older window to reach |
| Deleted content | A public index holds what was visible at collection time. Once removed, it cannot be re-collected |
The first row is the one that reshapes campaign measurement. A meaningful share of campaign delivery happens in formats that leave no public trace, so a freshness strategy that only addresses profile fields still misses them entirely.
How do you measure a vendor yourself?
With a controlled change and a stopwatch. This takes an afternoon and produces measured numbers rather than modelled ones, which is the only way to get a defensible figure for a specific vendor.
The method
- Use accounts you control, across each platform you care about, at a spread of sizes. You need to know the true value to measure the error.
- Make a detectable change and record the exact time. Post something, change the bio, or note the follower count from the native app at a fixed moment.
- Poll the vendor at a fixed interval until the change appears, recording the timestamp of each response.
- Compute the lag as first-detection time minus change time, per field and per platform.
- Separately, sample the index at random and compare each returned value against the live account. That gives you age of resting records rather than propagation lag, and the 2 measure different things.
- Report p50 and p95 per field per platform. A blended figure hides the field that is broken, which is the same error as a blended success rate.
# freshness_benchmark.py
# 2 measurements, often confused. Run both.
# 1. PROPAGATION LAG: how long a change takes to appear
# 2. RESTING AGE: how old a value is when you pull it cold
import time, statistics
def propagation_lag(vendor_call, platform, handle, field,
true_value, changed_at, poll_secs=300, max_hours=72):
"""Change something, then poll until the vendor reflects it."""
deadline = changed_at + max_hours*3600
while time.time() < deadline:
got = vendor_call(platform, handle).get(field)
if got == true_value:
return time.time() - changed_at # seconds
time.sleep(poll_secs)
return None # never propagated
def resting_age(vendor_call, platform, handle, field, live_value):
"""Compare a cold pull against the live account."""
r = vendor_call(platform, handle)
return {
"matches": r.get(field) == live_value,
"vendor_ts": r.get("data_fetched_at"), # ask for this field
"delta": abs((r.get(field) or 0) - (live_value or 0)),
}
def report(samples):
"""Report per field per platform. Never blend."""
ok = [s for s in samples if s is not None]
ok.sort()
if not ok: return {"p50": None, "p95": None, "never": len(samples)}
return {
"n": len(ok),
"p50": statistics.median(ok) / 3600, # hours
"p95": ok[int(len(ok)*0.95) - 1] / 3600,
"never": len(samples) - len(ok), # the real story
}
# The "never" count matters more than the percentiles. A change that
# has not propagated in 72 hours is not slow, it is a coverage gap
# that a latency figure would hide by excluding it from the sample.
That last comment is the part most benchmarks get wrong. If you compute a p95 only over the samples that eventually propagated, you have quietly excluded the worst results and produced a flattering number. Report the non-propagation count alongside the percentiles or the percentiles mislead.
What should you ask a vendor?
- What proportion of your index was refreshed in the last 30 days? The single most revealing question, because the honest answer is constrained by their cost structure.
- Do you expose a fetch timestamp per record? If not, you cannot tell fresh from stale and neither can they.
- Is the timestamp per record or per field group? A profile refreshed last week may carry post data from last quarter.
- How do you decide refresh priority? Demand-driven tiering is the right answer. A fixed schedule across a large index is not affordable.
- What is the propagation lag for a new post, per platform? Measured, with percentiles, not a marketing claim.
- Can I filter or sort by record age? If you can exclude stale records, you can trust the rest. If not, every result carries an unknown.
Where does Phyllo fit?
On the side where freshness is not a tiering decision. Phyllo's social data API reads from the platform through the creator's own authorisation, so a value arrives current at call time rather than being served from a snapshot with an unknown age. Change notifications are delivered rather than discovered by polling, which we compared in webhooks versus polling.
That is an architectural property, not a claim about engineering effort, and it only applies to creators who have connected. Per-platform field coverage is public at getphyllo.com/coverage and the API reference needs no sales call.
Where this does not help you. If you need data on creators who will never authorise your product, a public index is the only option and the staleness above is the price of admission rather than a vendor failing. Buy from a vendor who exposes a fetch timestamp, filter by age, and treat the long tail as indicative rather than as fact. That is the right way to use a public index, and it is a perfectly good way to run a discovery product.
The short version
Stop treating freshness as a vendor quality signal. The refresh arithmetic applies to everyone, and any large public index is mostly made of records that were last checked months ago. The variable is not whether a vendor has stale data, it is whether they will tell you which records are stale.
So buy on 2 things: a fetch timestamp per record, and the ability to filter by it. Everything else is a claim. And where a field genuinely has to be current, check whether it can be read live rather than indexed, because that is a different architecture rather than a better vendor.
Want values read live at call time rather than served from an index of unknown age? Get a demo
How stale is public social media data?
It depends on the refresh tier a record sits in. Because uniform refresh is unaffordable, a representative tier structure implies a median record age of about 160 days and a p95 of about 344 days.
Why can vendors not just refresh everything daily?
Cost. Refreshing 350 million profiles daily at $1 per 1,000 records is about $127,750,000 a year. Tiering the same index by demand costs roughly $1,040,000, which is why every large index tiers.
Is authenticated social data fresher than public data?
No meaningful age, which differs from being fresher. An authenticated read returns the current value at call time, so the only staleness is your cache. It applies only to accounts that authorised you.
Which social data fields go stale fastest?
Follower counts change continuously and post engagement moves fastest in the first 24 to 48 hours. Bios, display names and profile pictures are stable for months, so track a time-to-live per field.
How do I measure a social data vendor's freshness?
Change an account you control, log the time and poll the vendor until the change shows. Compare cold pulls with live values. Report per field and platform with percentiles and never-propagated counts.
What should I ask a vendor about freshness?
Ask what share of the index was refreshed in the last 30 days, whether a fetch timestamp is exposed per record or per field group, how refresh priority is set, and whether you can filter by age.
Can public data give me historical social metrics?
Not retrospectively. Public surfaces return current snapshots rather than time series, which is why vendor documentation advises starting to track a profile early so history accumulates forward.



