ChatGPT Citation Support Rate Holds at 5x the Sample
Machine Relations' live Fidelity cohort has now graded 8,258 ChatGPT citations at 53.4% supported, confirming the September frozen-sample rate of 54.8%. Perplexity, Gemini and Claude still have no live-cohort rate.
ChatGPT Citation Support Rate Holds at 5x the Sample
A frozen-corpus sample of 1,255 resolvable ChatGPT citations put its supported rate at 54.8%. A separate, growing live cohort has now graded 6,225 resolvable citations — five times the sample — at 53.4%. The number held. Three of the six Machine Relations Index engines still have no comparable live-cohort figure at all.
- Published: 2026-09-27
- Author: Paralax Editorial
- Canonical: https://paralax.ai/blog/chatgpt-citation-support-rate-holds-scale-2026
- Tags: ai-search, citations, measurement
A citation in an AI answer asserts something specific: that the linked page supports the sentence sitting beside it. That claim is testable, and the Machine Relations Index tests it nightly through its Answer-Source Fidelity instrument, grading whether a cited source actually backs the claim it anchors, cycle by cycle, across four of the six engines the Index observes.
On September 19, 2026, Paralax reported the instrument's first published read: on a frozen corpus of 10,066 citation events, ChatGPT's supported rate was 54.8% of 1,255 resolvable citations (77.0% resolvable coverage), and Perplexity's was 50.1% of 1,298 (47.6% coverage). Gemini and Claude published no rate at all — too few of their citations were resolvable to clear the instrument's evidence floor.
That frozen sample does not grow. It is a fixed corpus graded once, and every number built on it is permanent unless the corpus itself is re-collected. The instrument's second track does grow: a live, collection-time cohort that grades new citation events cycle by cycle as they are collected, currently 24 cycles deep, running September 2 to September 26, 2026.
What the live cohort says, read September 26, 2026
The live cohort has now graded 8,258 citation events, of which 6,225 were resolvable — a 75.4% resolvable coverage rate, close to the frozen sample's 77.0% for the same engine. Of those 6,225, 53.4% were graded Supported.
| Track | Engine | Citations sampled | Resolvable | Coverage | Supported, of resolvable |
|---|---|---|---|---|---|
| Frozen corpus, published 2026-09-19 | ChatGPT | 1,630 | 1,255 | 77.0% | 54.8% |
| Live cohort, read 2026-09-26 | ChatGPT | 8,258 | 6,225 | 75.4% | 53.4% |
| Frozen corpus, published 2026-09-19 | Perplexity | 2,726 | 1,298 | 47.6% | 50.1% |
| Live cohort, read 2026-09-26 | Perplexity | — | — | — | no rate published |
| Frozen corpus, published 2026-09-19 | Gemini | 4,349 | 392 | 9.0% | no rate published |
| Live cohort, read 2026-09-26 | Gemini | — | — | — | no rate published |
| Frozen corpus, published 2026-09-19 | Claude | 1,361 | 88 | 6.5% | no rate published |
| Live cohort, read 2026-09-26 | Claude | — | — | — | no rate published |
(Aggregate hash 652af19f, generated 2026-09-26T11:35:47Z; every rate is the instrument's own published supported_rate_over_resolved, read as-is, never recomputed.)
Two things follow, and they point in opposite directions.
ChatGPT's number is now durable, not a fluke of one sample. A single frozen read of 54.8% is a snapshot. The same rate, recovered independently from five times as many resolvable citations collected over three and a half weeks, landing at 53.4% — 1.4 points lower, well inside the range two draws from one underlying process should differ by — is a replication. For a brand or search team deciding how much to trust "ChatGPT cited us" as evidence that ChatGPT's answer actually reflects the page, the honest read is: roughly half of what ChatGPT cites checks out, and that number is not moving much as the sample grows.
Perplexity, Gemini and Claude have no comparable check yet. The live cohort currently reports resolvable citations only for ChatGPT; the other three engines carry no read in this track at the date above. That is a gap in what the instrument can currently confirm, not a claim that those engines' figures changed or held — the September 19 numbers for them stand only on the frozen sample, unconfirmed by the growing one. A team benchmarking engines against each other on citation reliability today can only make that comparison on a three-week-old, one-time snapshot for three of the four measured engines, and Google AI Mode and Google AI Overviews have never cleared the evidence floor on either track.
Why the distinction between the two tracks matters
The instrument's own publication rule is explicit about this, and it is worth stating in the reader's terms rather than the method note's: every rate is reported "of the N resolvable citations," never over the full citation count, and the two cohorts are never blended into one number. A frozen sample answers "what did the instrument find in one collection window." A live, accumulating cohort answers "does that finding keep showing up as more evidence arrives." Only the second is a test of durability, and only ChatGPT currently has enough live-cohort volume to run it.
This is not a Machine Relations peculiarity. It is the standard failure mode research on citation reliability keeps finding across the industry. Wallat et al. distinguish citation correctness — does the cited document support the claim — from citation faithfulness — did the model actually rely on that document rather than a prior belief it happened to match — and find up to 57% of citations in their tested systems fail the faithfulness test even when they pass a correctness check. The benchmark most attribution studies calibrate against, AttrScore (code, dataset, ACL version), exists precisely because early attribution-detection systems disagreed with human judgment often enough that a shared reference set was needed before anyone could compare rates across studies at all. Fidelity's own calibration certificate reports 91.2% agreement with a held-out reference set — a number worth knowing before trusting any rate the instrument publishes, including the ones in this table.
The pattern shows up outside the lab too. Columbia Journalism Review's Tow Center compared eight AI search engines on how accurately they credited the news articles they cited and found more than 60% of tested citations wrong or misleading across the field, a finding Nieman Lab's write-up of a follow-on Tow Center study on ChatGPT, Claude, Gemini and Grok ranked, again, worst for ChatGPT among the four. TechCrunch reported on the earliest of these Tow Center rounds when it first published. A separate, much larger domain — medical citation — got the same answer at a different scale: a Stanford-led audit published in Nature Communications built an automated framework to check whether LLM-cited medical references actually supported the claims attached to them, running the check across tens of thousands of claim-citation pairs. None of these studies measure the same engines, questions or time window as the Machine Relations Index, and none of their figures are interchangeable with Fidelity's — the value of citing them here is that the general finding, "citation presence and citation accuracy are different measurements, and the gap between them is large and well documented," keeps replicating across research groups, engines and subject domains, which is exactly why a brand team should not read "cited" as "correctly represented" on faith.
That gap is also becoming a live business risk that marketing teams are starting to name directly, not just researchers. A recent industry analysis on citation misrepresentation as brand risk argues that the GEO industry has spent its attention on being cited at all, when the more consequential failure mode is being cited and misquoted — a framing that lines up with what both the frozen and live Fidelity cohorts show for ChatGPT: roughly half of its resolvable citations do not actually support the claim they're attached to. A related read on the citation-manipulation backlash building around gaming AI search results makes the adjacent point that citation counts are already being treated as a target to game, which only works because most citation counts are being read uncritically as accuracy in the first place. Wikipedia's own description of retrieval-augmented generation is the plain-language version of the mechanism behind all of this: a model is grounded in retrieved documents specifically so its answer can be checked against them, and a check that nobody runs is not grounding, it is decoration.
What to do with this if you monitor citations
If a monitoring tool tells you your brand was "cited by ChatGPT," that is a presence measurement. It says nothing about whether the citing sentence reflects your page, and on the numbers above, close to half the time across two independent samples, it does not. Treat a ChatGPT citation as an unverified claim until you have read the sentence next to it against your own source. For Perplexity, Gemini and Claude, there is currently no comparably large, replicated per-engine rate to lean on at all — the September 19 read is the only one that exists, and it has not yet been checked twice.
Machine Relations' public release updates the aggregate daily at machinerelations.ai/measurement/answer-source-fidelity; the underlying study is at machinerelations.ai/research/citation-support-gap. Paralax will revisit this page's numbers once the live cohort clears the evidence floor for a second engine.
FAQ
Is ChatGPT's 53–55% supported rate good or bad? Neither figure is being compared here against a target; it is being compared against itself over time. What the two independent samples show is that the rate is stable, not that it is acceptable. Roughly half of ChatGPT's resolvable citations were graded as not supporting the claim they anchored, in both the frozen and the live measurement.
Why do Perplexity, Gemini and Claude have no live-cohort rate? The live, collection-time cohort has so far accumulated enough resolvable citation volume to report a rate for ChatGPT only; the other three engines carry no events in this track as of the September 26, 2026 read. That is a statement about what the instrument has been able to confirm so far, not a claim about those engines' underlying reliability.
What does "resolvable" mean, and why does it matter? Before a citation can be graded Supported or not, the instrument has to be able to fetch the cited page and isolate the specific claim the citation anchors. A citation that cannot be fetched, or cannot be tied to a single claim, is excluded from the rate entirely — it is neither counted as supported nor as failing. Coverage is the share of sampled citations that clear this bar; the supported rate only ever applies to that resolvable share, never to the full citation count.