Trace a statistic back to its source before you put it in a deck

Build a chain of sources from claim to primary data so you know what you're actually citing.

For anyone who cites numbers in public · 8 steps · 7 min

How it works today.

You spot a compelling statistic in a trade publication or LinkedIn post. The article links to another article, which references a report you can't download without filling out a lead form. You click through three more layers and find a consultant's blog that says the number came from 'industry research.' The deadline is this afternoon, so you screenshot the original article, paste the figure into your slide, and add the publication's name in eight-point type at the bottom.

Two weeks later someone asks where the number came from. You send the article link. They come back: the article has been updated and the figure is gone. You check the Internet Archive, find five different versions of the claim, and spend an hour reconstructing what you thought you knew. Next time you promise yourself you'll trace it properly. Next time you don't.

Before you start.

All of it has to be true, or step one fails in a way that is annoying to debug.

  • Access to a generalist agent (Claude, ChatGPT, or equivalent) that can browse the web or accept pasted URLs.
  • A reference manager, spreadsheet, or note file where you will record the chain of sources as you build it.
  • Read access to any paywalled databases or journals your organisation subscribes to, so the agent or you can retrieve PDFs behind login walls.
  • Willingness to mark a claim as unverified if the chain dead-ends before reaching primary data.

The steps.

  1. Paste the claim and its immediate source into the agent

    Copy the sentence containing the statistic and the URL or citation where you found it. Tell the agent you want to trace it to the original study, survey, or dataset. The agent will often extract the next link in the chain from the article text or footnotes. If the source is a screenshot or PDF, paste the relevant excerpt as text.

    Paste this
    I found this claim: [paste the sentence]. It appears here: [paste URL or citation]. Trace this statistic back through any intermediate sources until you reach the primary study, survey, dataset, or official report. List each source in the chain with its title, author or publisher, date, and URL or DOI.
  2. Check whether the agent can retrieve the next source in the chain

    The agent will return a link or citation to the next layer. Click it yourself or ask the agent to fetch the text. If the link is broken, the document is paywalled, or the agent cannot access it, note that in your tracker and try the Internet Archive or your institution's library portal. Do not let the agent guess what the source said.

  3. Ask the agent to extract the claim from the new source and compare it

    Once you have the text of the next source, ask the agent to find the sentence or table containing the statistic. Check whether the number, timeframe, or population has changed. If the second source says '23% of users' and your original article said '23% of companies,' record the discrepancy. The agent should quote the exact text, not paraphrase.

    Paste this
    Here is the text of [source title]: [paste excerpt or full text]. Find the sentence or table that contains the statistic about [topic]. Quote it exactly. Does this match the claim in the original article, or has the number, date range, or population changed?
  4. Repeat the extraction for each layer until you reach a primary source

    A primary source is a survey instrument, dataset, official register, or peer-reviewed study that collected the data. A press release about a survey is not primary. A consultancy's interpretation of a government report is not primary. Keep asking the agent to identify the next source and extract the claim until you reach something that describes how the data was gathered. Expect three to six layers.

  5. Ask the agent to summarise the methodology of the primary source

    Once you have the original study or dataset, ask the agent to extract the sample size, date range, geography, and method. If the primary source is a PDF, paste the methods section. If it is a dataset, ask for the documentation or readme file. Record this summary in your tracker. If the agent cannot find a methods section, the source may not be primary.

    Paste this
    Here is the primary source: [paste title and URL or text]. Extract the sample size, data collection period, geographic scope, and methodology. If this information is not present, say so.
  6. Check whether the primary source is publicly accessible

    Copy the URL or DOI of the primary source and open it in a private browser window. If it loads, add the link to your slide footer or reference list. If it requires a login, check whether your organisation has access or whether the authors have posted a preprint. If the source is not accessible to your audience, note that in your tracker and decide whether to use the claim.

  7. Record the full chain in your reference file

    List each source from the original article back to the primary study, with publication dates and URLs. Include any discrepancies you found—changed numbers, different populations, missing context. This chain is what you will send if someone asks where the figure came from. If the chain broke before reaching a primary source, write 'Source not verified' and do not use the statistic.

  8. Write the citation using the primary source, not the article

    In your deck or document, cite the original study or dataset, not the trade publication where you first saw it. If the primary source is hard to read, you can say 'According to [Study Name], as reported in [Article],' but the formal citation should point to the data. If you could not verify the chain, remove the claim or replace it with a verifiable one.

What you keep.

Automating the typing does not move the accountability. These stay with a person.

  • The decision whether to use a statistic whose chain you could not complete. An agent summarising a source is not the same as you confirming the number is real.
  • Judgment about whether a discrepancy between sources matters. If one source says 'users' and another says 'companies,' you decide whether that changes the claim's meaning.
  • The final citation and any explanatory footnote. You are accountable for what the audience believes after reading your slide.
  • Verification that the primary source's methodology is sound enough for your purpose. The agent can extract sample size and date range; you decide whether a survey of 80 people from 2019 supports a claim about the market today.

Once it works.

The first run is the demo. These are where the time actually comes back.

Run it on a loop

Every time you add a statistic to a slide library or fact sheet, run the chain and store it in a shared reference file so the next person does not re-trace the same path.

Fire it from an event

When a claim in your reference file is older than twelve months, re-run the chain to check whether the primary source has published an update or retracted the figure.

Scale it wider

If you cite more than five statistics in a single document, build the chain for all of them in one session and compile a sources appendix so reviewers can check your work without asking.

Hand off to another agent

When you send a draft to a subject-matter expert for review, include the source chain for each claim so they can verify the data without hunting through footnotes.

The work this replaces.

These are the O*NET work activities this workflow covers, and the categories of tool that address them.

Addressed by web, data and docs, generalist agents, models and routing