Consent and Data Retention for Authenticated Social Data

How to structure a consent record, handle revocation across platforms, set retention periods that survive backups, and keep the audit trail you must not delete.

Ronak Shah
Growth at Phyllo
September 26, 2026
3D clipboard with a green tick and signature beside a purple filing cabinet with a clock
Summarize this article with AI
GeminiChatGPTClaudePerplexityGrok

How to structure a consent record, handle revocation across platforms, set retention periods that survive backups, and keep the audit trail you must not delete.

This is some text inside of a div block.
  • A valid access token is not proof of consent; check consent state at the point of use, in the same transaction.
  • 3 events end a connection: token expiry, platform revocation and in-product withdrawal, each needing a different response.
  • Meta requires a Data Deletion Request Callback returning a confirmation code and a status URL the user can check.
  • Retention covers backups: a 12 month policy beside a 7 year backup rotation is the gap auditors look for; re-run the deletion index on restore.
  • Store granted scopes in the consent record, not requested ones, because TikTok and YouTube allow partial grants.

When a creator connects their social account to your product, you acquire 2 things: an access token, and an obligation. Most teams build carefully around the first and improvise around the second.

The improvisation usually shows up in the same 3 places. Consent is checked once at connection and never again. Retention is defined for the production database and forgotten for backups. And when a deletion request finally arrives, nobody can prove what was deleted or when.

This is a build guide for the second half: what a consent record should contain, how to detect and handle revocation across platforms that all signal it differently, how to set a retention period that survives a restore, and what you are obliged to keep. It is architecture guidance rather than legal advice, and your lawful basis and retention periods are decisions for counsel.

Why is an access token not proof of consent?

Because they describe different things at different moments. A token proves that consent existed when it was issued. It says nothing about whether consent exists now.

The gap between those 2 states is wider than it looks. A creator can revoke your app in their platform settings without telling you, and on most platforms you find out only when a call fails. Google goes further and revokes every refresh token carrying a YouTube scope when the user changes their account password, which is a revocation event with no intent behind it at all.

So the rule for anything that processes personal data: perform the consent check at the point of use, inside the same transaction as the work. A background job that was enqueued while consent was valid must re-check before it acts, because the queue may have been sitting for hours.

# Wrong: consent inferred from the message that enqueued the job
def process(job):
    fetch_and_store(job.account_id)   # consent checked when enqueued

# Right: check inside the worker, in the same transaction
def process(job):
    with db.transaction():
        state = consent_state(job.account_id, job.category)
        if state != "active":
            record_denial(job, state)     # useful metric, see below
            return
        fetch_and_store(job.account_id)

# Denied-after-enqueue events tell you how long work sits in the
# queue and where revocation races concentrate. Track the rate.

There is a real cost to this. A strongly consistent consent check on every protected operation costs latency and database capacity, and a long-lived cache lowers both while creating a window in which a withdrawn user is still served. For personal data that trade usually favours spending the read. If you choose to cache, write the decision down and put a bound on the window.

What should a consent record contain?

Enough to answer 4 questions later: who agreed, to what, when, and under which version of your terms. Anything less and the record cannot support a dispute.

FieldWhy it existsNotes
Subject identifierWho granted itUse a pseudonymous internal key rather than an email address wherever the record is read at volume
Platform and accountWhich connection this coversA creator with 4 connected accounts has 4 consent states, not 1
Scopes grantedWhat they actually agreed toStore what the platform returned, not what you requested. TikTok and YouTube both allow partial grants
Purpose or categoryWhat you may use it forSeparating purposes lets a user withdraw one without killing the whole connection
Consent versionWhich terms they sawIf your privacy notice changes materially, existing consent may not cover the new purpose
Granted timestampWhenWith timezone. This is the field disputes turn on
Stateactive, withdrawn, expired3 states, not a boolean. Withdrawn and expired need different handling
State change historyAppend-only transitionsWho changed it, which category, when, and a correlation ID for the request

The scopes row is the one most often got wrong. On TikTok and YouTube, a user can approve part of your request, so the scope string you sent is not a record of anything. On LinkedIn the flow fails entirely if any scope is unavailable, and on Instagram the grant follows the permissions your app was approved for. 4 platforms, 4 behaviours, 1 field that has to hold the truth for all of them.

If you would rather not maintain consent state across 4 platform models, we do it as part of the connection layer. See how it works

What are the 3 ways a connection ends?

They look similar in your logs and they need different responses. Conflating them is why users get told to reconnect when a refresh job would have fixed it, and why revoked accounts keep getting retried forever.

EventWhat happenedCorrect response
Token expiryThe access token reached its lifetime. The refresh token is still validRefresh silently. The user should never know
Platform revocationThe creator removed your app in their platform settings, or something external invalidated the grantStop reading. Mark the connection dead in your interface. Prompt a reconnect
In-product withdrawalThe creator withdrew consent inside your product, or deleted their account with youStop reading, revoke the token at the platform, and run your deletion process

The third row carries an obligation the first 2 do not. If a user withdraws inside your product, revoking the token at the platform is your job, not theirs. Leaving a live token on a connection the user believes they severed is the kind of detail that surfaces badly in a security review.

How does Meta's deletion callback work?

Meta requires apps that access Facebook user data to implement a Data Deletion Request Callback, and to state in the privacy policy how users can request deletion. This is a platform terms requirement, separate from whatever privacy law applies to you.

The user triggers it from their Facebook profile, under Settings and Privacy, then Settings, then Apps and Websites, using the Send Request button. Meta then calls your endpoint.

# Meta POSTs a signed_request to the URL you configure in
# App Dashboard > Settings > Data Deletion Request URL

def data_deletion_callback(request):
    signed = request.form["signed_request"]

    payload = verify_and_decode(signed, APP_SECRET)  # validate first
    user_id = payload["user_id"]                      # app-scoped ID

    subject = resolve_subject(user_id)
    ticket  = create_deletion_ticket(subject)         # track it
    enqueue_deletion(ticket.id)                       # do the work async

    # Meta expects BOTH fields back
    return {
        "url": f"https://yourapp.com/deletion-status/{ticket.code}",
        "confirmation_code": ticket.code,
    }

# The url must show the user the status of their request.
# A page that always says "complete" is not a status page.

3 things to get right here. Validate the signature before you trust anything in the payload. Do the deletion asynchronously, because the callback needs a fast response and deletion across your systems is not fast. And make the status URL real, since its entire purpose is letting the user verify that something happened.

Note the identifier. Meta sends an app-scoped user ID, which means you can only resolve it if you stored that ID at connection time. If your schema keys connections on an internal ID and never persisted the platform identifier, you cannot honour the callback at all.

How long should you keep the data?

For as long as the purpose requires, and no longer. Under GDPR the storage limitation principle is that personal data should be kept in identifiable form for no longer than is necessary for the purposes for which it is processed. Once the purpose is fulfilled, the data should be deleted or anonymised.

That is deliberately not a number, which is inconvenient and correct. A retention period is a decision about purpose, so it has to be made per data category rather than globally. A worked example of how the categories differ:

Data categoryTypical driverWhat makes it expire
Access and refresh tokensOperationalRevocation, or the platform lifetime. There is no reason to hold a dead token
Profile snapshotsProduct purposeWhether your product still needs to show the creator their own history
Performance metricsProduct purposeOften longer, because trend reporting is the purpose
Earnings dataPurpose plus regulationMay carry statutory retention obligations that outlive the user relationship
Consent recordsEvidentialOutlives the data it authorised. See the section below
Audit logsEvidentialOutlives the data. Append-only and access-restricted

Whatever you decide, write the periods down per category in your record of processing activities. A retention schedule that exists only in someone's head cannot be shown to anyone, and the point of the schedule is that it can be shown.

Why do backups break most retention policies?

Because deletion from a live database does not reach a snapshot taken last month, and storage limitation applies to backups as well as production systems.

This produces a specific and common gap: a documented retention period of 12 months, a backup rotation that keeps 7 years, and no process connecting the 2. If your backup retention contradicts the deletion periods in your documentation, that contradiction is exactly what an auditor is looking for.

You cannot practically delete individual records from immutable backups, and you are not expected to. The workable pattern is a deletion index plus a mandatory post-restore step.

# 1. When a deletion is honoured, record a NON-IDENTIFIABLE marker.
#    Never key this index on email, name or platform handle, or the
#    index becomes a second repository of personal data.

deletion_index.insert(
    record_ref = internal_row_id,      # or a salted internal hash
    category   = "profile_snapshot",
    deleted_at = now(),
)

# 2. Make re-deletion part of the restore runbook, not an afterthought.

def restore(backup):
    load(backup)
    for entry in deletion_index.all():        # MANDATORY step
        purge(entry.record_ref, entry.category)
    mark_restore_complete()

# A restore that skips step 2 silently resurrects deleted records.

Document that runbook in your data protection materials. The pattern only counts if it is written down and actually followed, and the step is easy to skip under the pressure of an incident, which is precisely when restores happen.

What must you not delete?

The evidence that you honoured the deletion. This sounds contradictory and it is not: erasing personal data and destroying your record of having erased it are different acts, and only the first is required.

If you delete a user's data and also delete every trace that the request existed, you have removed your own ability to demonstrate compliance. The audit trail should be append-only, access-restricted, and written in a form that is not itself a store of personal data.

KeepWhy
The consent record and its state transitionsIt evidences what was agreed and when it ended
A deletion ticket with a timestamp and a scopeIt evidences that the request was received and fulfilled
The non-identifiable deletion indexIt makes post-restore re-deletion possible
Denial and access-decision logsThey show the gate was working

What those records should contain: a pseudonymous subject key, the category, the decision, the consent version, the actor type, a correlation ID and a latency figure. What they should not contain: tokens, raw personal data, or anything that would make the log a replacement copy of what you deleted. Log state transitions, not secrets.

What does the architecture look like end to end?

  1. At connection. Write the consent record with the granted scopes, the platform account identifier, the purpose and the consent version. Store the platform-scoped user ID, because a deletion callback will send it back to you.
  2. At every use. Check consent state in the same transaction as the work. Do not infer it from a queue message or a valid token.
  3. On failure. Classify the error into expiry, platform revocation or withdrawal. Refresh, prompt, or delete accordingly.
  4. On withdrawal. Revoke the token at the platform, transition the consent record, enqueue deletion by category, and give the user something they can check.
  5. During deletion. Work one bounded category per transaction and write a completion marker for each. If a worker crashes after the commit but before acknowledgement, redelivery is harmless because the marker has already committed.
  6. After restore. Re-run the deletion index. Every time, as part of the runbook.
  7. Continuously. Alert on the refresh failure rate rather than individual failures, and on denied-after-enqueue volume, which reveals where revocation races concentrate.

What are the most common mistakes?

  • Treating a valid token as proof of consent. The single most common one, and the hardest to see, because nothing fails.
  • Storing requested scopes instead of granted scopes. The 2 diverge on any platform that allows partial grants.
  • A boolean consent flag. Withdrawn and expired need different handling, so you need at least 3 states.
  • One consent state per user rather than per connection. A creator with 4 accounts can withdraw 1.
  • Retention defined only for production. Backups are covered by the same principle.
  • A deletion index keyed on personal data. It turns your compliance tooling into another copy of what you deleted.
  • Deleting the audit trail along with the data. You have then removed your ability to prove you complied.
  • Not persisting the platform-scoped user ID. Meta's deletion callback sends it, and without it you cannot resolve who to delete.

Where does Phyllo fit?

We hold this layer for the connections that run through us. When a creator connects an account through Phyllo's social data API, the consent record, the granted scopes, the token lifecycle and the revocation handling sit with us, and your product reads normalised data through one schema across 25+ platforms rather than tracking 4 different revocation behaviours itself.

That is the practical argument for a connection layer on this topic specifically. The consent model is not hard on 1 platform. It is hard because Instagram, TikTok, YouTube and LinkedIn each signal revocation differently, expire tokens on different schedules, and treat partial grants differently, and your compliance posture is only as good as the weakest of the 4 implementations. Our security and compliance position is at getphyllo.com/security, and identity resolution keeps a creator's connected accounts tied to 1 record so consent state is per connection rather than guessed.

Where we are not the answer: your own retention periods, your lawful basis and your privacy notice are yours. We hold the connection, not your policy. If you are screening people who never authorised anything, consent is a legal process rather than a technical one, and we covered that in what social screening is.

The short version

Write the consent record at connection with the scopes you were actually granted, then check the state at every point of use rather than trusting the token. Classify connection failures into expiry, platform revocation and withdrawal, because all 3 look alike and need different handling.

Set retention per data category, apply it to backups through a deletion index and a post-restore step, and keep the evidence that you honoured deletions even after the data itself is gone. None of it is complicated. It is just rarely built until somebody asks for it in writing, and by then the gap is a year wide.

Want the consent and revocation layer handled across every platform you connect? Get a demo

Is an OAuth token proof that consent still exists?

No. A token proves consent existed when issued. Users can revoke an app in platform settings unseen, and Google revokes YouTube-scoped refresh tokens on a password change. Check consent state at use.

What should a consent record contain?

A subject identifier, platform and account, scopes actually granted, purpose, consent version, granted timestamp, a state of active, withdrawn or expired, and an append-only history of state changes.

How long can I keep authenticated social data?

For as long as the purpose requires. GDPR storage limitation: keep personal data in identifiable form no longer than necessary, then delete or anonymise. Set periods per category with counsel.

Does data retention apply to backups?

Yes. Storage limitation covers backups too, so a 12 month retention period beside a 7 year backup rotation is a gap. Keep a non-identifiable deletion index and re-apply it after every restore.

What is the Meta data deletion callback?

Required for apps using Facebook user data. Meta sends a signed POST with an app-scoped user ID to your configured URL; verify the signature, start deletion, return a confirmation code and status URL.

Do I have to delete the audit trail too?

No, and generally you should not. Erasing personal data and destroying evidence you erased it are different acts. Keep an append-only, pseudonymous record that the request was received and fulfilled.

What happens if a creator revokes on the platform rather than in my product?

You usually get no notification and find out when calls fail. Classify that error as revocation, not expiry: stop reading, mark the connection dead in your interface, and prompt a reconnect.

Table of Content
See Phyllo in action
  • No Credit card required
  • GDPR and SOC Compliant
  • 30-min Onboarding
Book a Demo →

Be the first to get insights and updates from Phyllo. Subscribe to our blog.

Ready to get started?

Sign up to get API keys or request us for a demo