How to Build an Influencer Discovery Product: Reference Architecture

A reference architecture for influencer discovery: index design, identity deduplication, tiered refresh economics, and where the verification handoff belongs.

Ronak Shah
Growth at Phyllo
September 28, 2026
3D magnifying glass over a stack of profile cards beside tiered discs wrapped in a refresh arrow
Summarize this article with AI
GeminiChatGPTClaudePerplexityGrok

A reference architecture for influencer discovery: index design, identity deduplication, tiered refresh economics, and where the verification handoff belongs.

This is some text inside of a div block.
  • Refreshing 350 million profiles in an influencer discovery index daily at $1 per 1,000 records costs about $127.75 million a year.
  • Demand-tiered refresh of the same index costs about $1.04 million, roughly 123 times less, making tiering the only viable design.
  • Deduplication decides the headline count: a creator on 5 platforms is 1 record or 5, so published counts are not comparable.
  • Buyers filter on audience demographics, which a public index models rather than measures, so derivation quality caps the product.
  • Verification is the second half: after discovery shortlists about 20 candidates, they connect and measured data replaces estimates.

An influencer discovery product looks like a search problem and is really a freshness problem. Building the index is a solved exercise. Keeping several hundred million records current enough to search is where the money goes and where the design decisions actually live.

The second thing worth establishing early: discovery runs on public data. Creators who have never heard of your product will not authorise it, so everything in the index is either publicly visible or modelled from what is publicly visible. That constrains which fields can exist at all, and it is covered in authenticated versus public social data.

What follows is the 8-stage architecture, the refresh arithmetic that determines whether the product is viable, and where the handoff to verified data belongs.

What are the pipeline stages?

StageWhat it doesWhat it hands on
1. AcquisitionCollects public profile and content data at volumeRaw per-account records with a fetch timestamp
2. Identity and dedupResolves accounts to creators and merges duplicatesOne creator record with N linked accounts
3. DerivationComputes engagement rate, growth, category, languageDerived fields with their formula version
4. IndexingWrites searchable documents with freshness metadataA queryable index
5. Query and rankingFilters, sorts and scores resultsA result set with per-field provenance
6. RefreshRe-collects on a tiered schedule driven by demandUpdated records, and an age distribution
7. VerificationReplaces estimates with measured data once a creator connectsA shortlist with real numbers
8. ServingDelivers results through UI or API with cachingThe product

Stages 2 and 6 are the ones that decide whether the product works. Everything else is standard engineering.

Why is identity deduplication the hardest stage?

Because it determines both your data quality and the headline number you publish, and the 2 pull in opposite directions.

A creator with accounts on Instagram, TikTok, YouTube, X and a newsletter is 1 person. Indexed naively they are 5 records. Deduplicate properly and your published profile count drops by a factor that makes your index look smaller than a competitor who did not bother.

That is a real commercial incentive not to deduplicate, and it explains why published counts across this market are not comparable. The same vendor is described with different totals across different sources, and part of that spread is definitional rather than factual. We audited the pattern in the universal API guide.

4 signals are worth combining, in roughly this order of reliability.

  1. Explicit cross-links. A link-in-bio page, or a profile that names the handle on another platform. Strong evidence.
  2. Verified connections. If a creator has connected accounts through any product, that link is a fact rather than an inference.
  3. Handle and name similarity. Cheap, noisy, and produces false merges on common names.
  4. Content fingerprinting. The same image or video posted across platforms. Expensive, and strong when it hits.

Whatever you combine, carry a confidence score on every link and never merge silently below a threshold. A wrongly merged creator record pollutes every metric derived from it, and unpicking it later is far harder than not doing it.

How should derived fields be computed?

Deterministically, versioned, and with the formula recorded alongside the value. Derived fields are where discovery products diverge most from each other, and almost none of the divergence is about data quality.

Engagement rate is the clearest case. There is no standard denominator, and each defensible choice produces a different number from identical data.

DenominatorWhat it measuresWhy someone picks it
FollowersEngagement against audience sizeThe most common. Penalises large accounts structurally
Reach or impressionsEngagement against actual viewsThe most accurate, and unavailable in public data
Recent post averageEngagement against the creator's own baselineBetter for comparing a creator with themselves over time
Median rather than meanSame, resistant to one viral postProduces a materially different ranking on creators with a spike

Pick one, publish which one, and version it. When you change the formula, every historical value computed under the old one becomes incomparable, so the version belongs in the record rather than in a changelog.

The same discipline applies to anything you cannot measure. Audience demographics in a public index are modelled, typically by sampling visible followers and projecting the distribution. That is legitimate and useful, and it becomes a problem only when a modelled number is presented as a measured one. Mark every derived field as estimated in the data model so the label survives all the way to the interface.

# Every derived field carries its method and version.

derived = {
    "engagement_rate": {
        "value":     0.0412,
        "method":    "mean_engagement_over_followers",
        "version":   "er-v3",
        "basis":     "last_12_posts",
        "measured":  False,          # renders as "estimated" in the UI
    },
    "audience_country": {
        "value":     {"US": 0.41, "GB": 0.12, "CA": 0.08},
        "method":    "follower_sample_projection",
        "version":   "aud-v2",
        "sample_n":  2500,
        "measured":  False,
    },
}

# measured: False is not a disclaimer. It is a field the UI reads
# and the API returns. A brand that challenges a number needs to
# know whether it was counted or modelled, and so does your support team.

If you want to know which fields can be measured rather than modelled, our per-platform coverage list is public. See the coverage list

What actually determines refresh strategy?

Arithmetic. Work the numbers once and the design decides itself.

Take an index of 350 million profiles, which is the scale the larger discovery products advertise, and a collection cost of $1 per 1,000 records, which is the low end of published scraper API rates. One complete pass costs $350,000.

Refresh cadenceAnnual costVerdict
DailyAbout $127,750,000Not a business
WeeklyAbout $18,200,000Not a business
MonthlyAbout $4,200,000Exceeds most companies' entire infrastructure budget
QuarterlyAbout $1,400,000Affordable, and the data is a quarter stale

So uniform refresh at any useful cadence is impossible. The alternative is to refresh by demand, because the distribution of what people actually search is extremely uneven. Most of an index is never returned in a single result.

TierShare of indexCadenceAnnual cost
Hot, appears in results often0.1%DailyAbout $127,750
Warm, appears occasionally1%WeeklyAbout $182,000
Cool, rarely queried10%MonthlyAbout $420,000
Cold, the long tail88.9%AnnualAbout $311,150

Total: roughly $1.04 million a year, against $127.75 million for uniform daily. About 123 times cheaper, for an index where the records people actually look at are fresher than uniform daily would have made them, because the hot tier gets the same treatment while the tail stops consuming budget.

That result reframes the whole product. Tiering is not a cost optimisation applied to a working design. It is the only design, and everything else follows from it.

# Demand-driven tiering. Promotion beats a static schedule.

def tier_for(profile):
    # Signals that a profile matters, strongest first
    if profile.in_active_shortlist:   return "hot"     # someone is deciding
    if profile.appeared_in_results_7d: return "warm"
    if profile.appeared_in_results_90d: return "cool"
    return "cold"

CADENCE = {"hot": 1, "warm": 7, "cool": 30, "cold": 365}   # days

def due_for_refresh(profile, now):
    tier = tier_for(profile)
    return (now - profile.fetched_at).days >= CADENCE[tier]

# A profile promotes itself by being searched. That is the whole
# mechanism, and it is better than any category-based rule,
# because it tracks what your customers actually care about today.

2 consequences to design for. A profile entering a shortlist should trigger an immediate refresh rather than waiting for its tier, because that is the moment somebody is making a decision on the number. And a cold profile returned in a search should be served with its age visible, since a 9 month old follower count presented without a date is worse than no follower count.

Why is freshness a query-time concern?

Because in a tiered index every record has a different age, and hiding that turns a reasonable engineering trade into a data quality claim you cannot support.

Every document in the index should carry its fetch timestamp per field group, and every result should return it. Then the interface can show it, filters can use it, and ranking can prefer fresher records when 2 candidates are otherwise equal.

A filter worth building: let users exclude records older than a threshold. It looks like a feature that shrinks your result set and makes the index look smaller. It is actually the feature that makes the results trustworthy, and the buyers who ask for it are the ones evaluating you most carefully.

What should the query layer optimise for?

Filters, not corpus size. Every vendor in this market advertises how many creators they hold, and buyers search on audience location, audience age, engagement rate and category.

That mismatch has a design consequence. The fields people filter on hardest are the modelled ones, so your filter quality is limited by your derivation quality rather than by your index size. Doubling the corpus does nothing for a user filtering on audience country if the country projection is weak.

Practical priorities for the query layer:

  • Facet on the filters people actually use, and instrument which filters are applied so the list is evidence rather than opinion.
  • Return provenance per field, so the interface can distinguish measured values from modelled ones without a second call.
  • Support exclusion as a first-class operation. Brands more often need to rule creators out than find new ones.
  • Make ranking explainable. A user who cannot tell why a creator ranked third will not trust the first 2 either.

Where does verification fit?

At the shortlist boundary, and it is the second half of the product rather than an add-on.

Discovery narrows a very large index to perhaps 20 candidates using public and modelled data. That is the right tool for that job. But the decision that follows, committing budget to a creator, is made on numbers the public index cannot produce: real reach, real audience breakdown, performance on formats that expire, and earnings.

So the architecture has a handoff. A shortlisted creator is invited to connect their accounts, and from that point their record carries measured values alongside the estimates. 3 rules keep that clean.

  1. Measured always beats modelled on the same field. Never average them, and do not prefer the newer value across sources. A measured reach from yesterday beats a modelled reach from an hour ago.
  2. Keep both values. The estimate is still useful, because comparing it to the measurement is how you tune your model.
  3. Show which is which. A brand asked to approve a budget deserves to know whether the reach figure was counted or inferred.

This is also where the economics change. Public collection bills per record, so refreshing the same profile repeatedly multiplies your cost. Connected access bills per account, so re-reading is close to free, which makes the shortlist exactly the population where connection pays for itself. The cost models are compared in social data API pricing.

Where does throughput break?

Not at query time, and not at indexing. Modern search infrastructure handles this corpus size without difficulty. The constraint is collection throughput against vendor rate limits, and it binds on refresh rather than on search.

# Refresh capacity, and whether your tiers actually fit.

def daily_refresh_capacity(records_per_day_budget):
    return records_per_day_budget

def tier_demand(index_size, tiers):
    """tiers: {name: (share, cadence_days)}"""
    return sum(index_size*share/cadence for share, cadence in tiers.values())

TIERS = {"hot":  (0.001, 1),
         "warm": (0.01,  7),
         "cool": (0.10,  30),
         "cold": (0.889, 365)}

demand = tier_demand(350_000_000, TIERS)
# If demand exceeds your daily budget, the cold tier stretches first.
# Never let the hot tier degrade: it is the only one anyone sees.

Design the degradation explicitly. When collection capacity drops, whether through a rate limit change or a provider outage, the cold tier should stretch and the hot tier should hold. A system that degrades uniformly makes the records people are looking at worse in order to keep updating records nobody has queried in a year.

What should you build and what should you buy?

StageDefaultReasoning
AcquisitionBuyPer-platform coverage and access changes are a permanent maintenance load, and rates are published
Identity and dedupBuy or hybridAccuracy comes from cross-platform signals you may not hold. This is the stage where quality is hardest to reach alone
DerivationBuildYour formulas are your product's point of view
IndexingBuildStandard search infrastructure
Query and rankingBuildThis is what customers are actually buying
Refresh orchestrationBuildIt encodes your economics and nobody else can tune it for you
VerificationBuyConsented access across platforms means app review and token lifecycle on each one
ServingBuildYour product

What are the most common mistakes?

  • Uniform refresh. The arithmetic above rules it out before any other consideration.
  • Hiding record age. In a tiered index every record has a different age, and concealing it converts an honest trade-off into a quality claim you cannot defend.
  • Optimising corpus size. Buyers filter on modelled fields, so derivation quality caps the product long before index size does.
  • Silent identity merges. A wrongly merged creator pollutes every derived metric downstream.
  • Unversioned derived formulas. Change the engagement rate calculation and every historical value becomes incomparable.
  • Presenting modelled values as measured. Fine until a brand challenges a reach figure in a campaign post-mortem.
  • Degrading uniformly under pressure. Stretch the cold tier, protect the hot one.
  • Treating verification as a later phase. It is the half of the product where the money is committed.

Where does Phyllo fit?

Not in the index. We are a consented layer, and discovery by definition operates on creators who have not authorised anything, so a public index is the correct tool for stages 1 and 4 and we are not a substitute for it.

We fit in the 2 stages either side of it. Identity resolution and account linkage help stage 2, because verified connections are the strongest cross-platform signal available and inference is what everyone else is working from. And Phyllo's social data API is the verification handoff at stage 7: a shortlisted creator connects, and their estimated reach and modelled demographics are replaced with measured engagement and real audience data across 25+ platforms through one integration. Influencer vetting covers the brand safety check on the same shortlist.

Per-platform field coverage is public at getphyllo.com/coverage and the API reference needs no sales call, so you can check which of your fields can be measured before deciding where the handoff belongs.

Where we are the wrong purchase, stated plainly. If what you need is breadth across creators who will never connect, buy a public index. Modash, EnsembleData and the web data platforms exist for that and they do it better than any consented layer can, because the job requires exactly what consent cannot provide. We become relevant at the moment your user stops browsing and starts deciding.

The short version

Do the refresh arithmetic before you design anything else, because it eliminates most of the options. Uniform refresh at any useful cadence is unaffordable, demand-driven tiering is roughly 2 orders of magnitude cheaper, and every architectural decision after that follows from the tiering.

Then be honest in the data model about what you measured and what you modelled, carry record age into the results, and put the verification handoff at the shortlist rather than at the end. Discovery gets a brand to 20 candidates. It should not be the thing that gets them to a signature.

Want verified data on the shortlist your discovery product produces? Get a demo

How often should an influencer index be refreshed?

By tier, not uniformly. Refreshing 350 million profiles daily at $1 per 1,000 records costs about $127.75 million a year. Tiering costs about $1.04 million, roughly 123 times less.

How do you decide which profiles to refresh first?

By demand. Profiles in an active shortlist refresh immediately, profiles seen in recent results refresh often, and the long tail rarely, which tracks what customers care about today.

Why do influencer platforms report different profile counts?

Partly deduplication. A creator with accounts on 5 platforms is 1 record or 5 depending on a design decision, so a platform that deduplicates publishes a smaller number than one that does not.

How should engagement rate be calculated?

No standard exists. Followers is the most common denominator, reach the most accurate and not public, and median rather than mean reranks creators with a viral post. Pick one, publish it, version it.

Can a discovery index include audience demographics?

Only as estimates. Platforms release audience age, gender and location to the account holder alone, so a public index models them, usually by sampling visible followers. Label them as estimated.

Where does verified data fit in a discovery product?

At the shortlist boundary. Shortlisted creators connect their accounts before budget is committed, and a measured value always beats a modelled one on the same field. Keep both to tune the model.

What limits throughput in a discovery product?

Collection against vendor rate limits, which binds on refresh rather than search. When collection capacity drops, the cold tier should stretch and the hot tier should hold.

Table of Content
See Phyllo in action
  • No Credit card required
  • GDPR and SOC Compliant
  • 30-min Onboarding
Book a Demo →

Be the first to get insights and updates from Phyllo. Subscribe to our blog.

Ready to get started?

Sign up to get API keys or request us for a demo