ChatGPT 54.8%, Perplexity 50.1%: The Share of Resolvable AI Citations That Support Their Claim
A calibrated instrument graded 15,525 citation events across four AI answer engines. Of 1,255 resolvable ChatGPT citations, 54.8% supported the claim they anchored, at 77% resolvable coverage; of 1,298 Perplexity citations, 50.1%, at 47.6% coverage. Gemini and Claude publish no rate at all.
A citation in an AI answer is an evidence claim. It asserts that the linked page supports the sentence it sits beside. That assertion is testable, and almost nobody tests it — because testing it means pulling the cited page, isolating the specific claim the citation anchors, and grading the pair.
Machine Relations, the public research initiative that publishes the Machine Relations Index, has been running that test nightly since September 2. Its citation support gap study and the live aggregate behind it now cover 15,525 graded citation events across four answer engines.
The headline result is per-engine, and it is not close to the confidence a citation implies.
The per-engine numbers
Read from the published aggregate on September 19, 2026 (aggregate hash 90f6711dd14e473b, generated 11:34:53 UTC), frozen-corpus cohort:
| Engine | Citations sampled | Resolvable | Resolvable coverage | Supported, of resolvable |
|---|---|---|---|---|
| ChatGPT | 1,630 | 1,255 | 77.0% | 54.8% |
| Perplexity | 2,726 | 1,298 | 47.6% | 50.1% |
| Gemini | 4,349 | 392 | 9.0% | no rate published |
| Claude | 1,361 | 88 | 6.5% | no rate published |
Every rate in that table is a rate over the resolvable set, not over all citations. That is the instrument's own publication rule, and it is worth stating plainly: the correct sentence is "of the 1,255 resolvable ChatGPT citations, 54.8% were supported, and resolvable coverage was 77%." There is no defensible blended number across engines, and the study refuses to print one.
What "resolvable" hides, and why it is the more interesting column
Before you can ask whether a citation supports a claim, two things have to be true. You have to be able to tie the citation to a discrete claim rather than to a whole paragraph of synthesis. And you have to be able to read the cited page.
In the frozen corpus of 10,066 events, 6,017 — 59.8% — were unmappable: the answer's shape gave no discrete claim for the citation to anchor. Another 666 pages were unreadable and 350 graded uncertain. Only 3,033 survived to a grade.
That is why coverage spreads roughly twelvefold across engines, from ChatGPT's 77% to Claude's 6.5%. Whether a citation can be pinned to a claim at all is an engine property — a function of how that engine formats answers — before it is ever a property of the cited page. An engine that writes in discrete, individually cited assertions exposes itself to this test. An engine that writes flowing prose with a source list at the end does not.
Two fetch bases, one answer
The obvious objection to the corpus number is timing. Its cited pages were fetched at grade time, days after each answer was produced, so a page that changed in between would grade against text the engine never saw.
The study's second cohort removes that objection. Since cycle 0001 on September 2, every new event's cited page is snapshotted at collection time — minutes after the answer is produced — and grading consumes only that snapshot. That series now holds 5,459 events. Of its 4,013 resolvable ChatGPT citations, 52.0% were supported, at 73.5% resolvable coverage.
54.8% at grade time. 52.0% at collection time. Different denominators, different fetch bases, separate cohorts that the study never merges — and the same answer to within three points. The gap is not an artifact of when the page was read.
The collection-time series also isolates a failure mode the corpus buried. Its unmappable count is zero, but 1,063 of its 5,459 events — 19.5% — were unreadable: the cited page could not be fetched back at all, minutes after an engine had cited it. A further 383, 7.0%, ended uncertain.
Two engines that publish nothing
Gemini and Claude sit below the study's evidence floor and publish no support rate. Gemini resolved 392 of 4,349 citations; Claude, 88 of 1,361.
That is the correct behaviour and it is rarer than it should be. A 9% coverage sample will produce a number, and the number will be confidently wrong, because the 9% that resolved is not a random draw from the 100% that did not. Withholding is the honest move, and it is the reason the two rates that are published can be read as measurements rather than marketing.
Both engines added or changed web retrieval during the observation window — Anthropic's web search tool reached general availability in that period, and Google's AI features documentation describes the indexing and snippet eligibility conditions that govern whether a page can appear as a supporting link in AI Overviews and AI Mode at all. Coverage that thin is as likely to be an artifact of answer format and retrieval policy as of anything a publisher did.
How the instrument is calibrated
Every citation event is graded by two models from different families, and a blind third-family judge is called only when they disagree. If the judge matches neither candidate, or any grader fails to return a label, the item is left unresolved and recorded as uncertain — never forced to a side. Control fixtures re-seed every run; one control miss halts the run. Across the published calibration, control accuracy was 30 of 30 and inter-model agreement was κ 0.647, raw agreement 72.2%.
Against external human labels the instrument scores 83 of 91 items correct, 91.2%, with a Wilson lower bound of 83.6%. Against AttrScore, the public attribution-evaluation benchmark from the EMNLP 2023 Findings paper Automatic Evaluation of Attribution by Large Language Models (preprint, code), pinned at dataset revision 467dcdd2, binary agreement is 55 of 60, 91.7%, Wilson lower bound 81.9%.
Where it is weak, stated by the instrument
The same benchmark record publishes the failure. AttrScore labels three classes, and agreement is not even across them: 19 of 22 attributable (86.4%), 19 of 20 not attributable (95.0%), and 6 of 18 extrapolatory (33.3%). Eleven of the eighteen extrapolatory items were graded not attributable.
Extrapolatory is the hard middle — the cited page is about the right thing, and the claim reaches past what it actually says. The instrument reads most of that middle as unsupported. That is a real bias with a real direction, and it is why publication eligibility rides on the binary supported/unsupported metric and never on the three-way subtype breakdown, which the record marks diagnostic-only.
Cumulatively the instrument has spent $53.87 across 20,509 model calls to grade 7,046 events out of 15,525. Reproducibility is the point of those receipts: the benchmark revision is pinned, row-level evidence is content-hashed, and the aggregate carries its own hash.
What a publisher can actually change
Two of the three failure modes are not yours. Whether an engine writes discrete cited claims is the engine's choice. Whether a given claim reaches past its source is the model's.
The third one is yours, and it is 19.5% of the collection-time series: the cited page could not be read back. A page that cannot be fetched cannot be verified by anyone — not by this instrument, not by a buyer checking a vendor claim, not by the next retrieval pass. The relevant controls are documented and boring: RFC 9309 governs robots exclusion, and the engines publish their own agents — OpenAI's bots, Google's crawlers, Perplexity's, Anthropic's. Crawler identity is not self-certifying either; Cloudflare has documented undeclared crawlers evading no-crawl directives, which is a reason to verify by reverse DNS rather than by user-agent string.
The second lever is the claim itself. Support is graded against the specific assertion a citation anchors, not against the page's general topic. A page that states its claim, its denominator and its date in checkable sentences can be graded supported. A page that gestures at the claim cannot — and in a retrieval-augmented answer, gesturing is what gets read as extrapolation.
What this does not say
It does not say half of AI answers are wrong. Support is a property of the (claim, source) pair, not of the answer: an unsupported citation often sits beside a true sentence that the linked page simply does not prove.
It does not rank the engines on accuracy. ChatGPT scores highest here partly because its answer format is the most gradeable, which is a different virtue from being right.
And it is not a trend. Two cohorts agreeing at one reading is a convergence, not a direction. The collection-time series is 17 cycles old; the first honest trend statement will need more of them.
FAQ
What does "supported" mean here? That the cited page contains text supporting the specific claim the citation anchors, judged by two independent graders with a blind judge on disagreement, under the five-label set Supported, Amplified, Contradicted, Misattributed, Fabricated, reported as a binary supported rate.
Why is there no single number across all four engines? Because coverage differs by roughly twelvefold. A pooled rate would be dominated by whichever engine happened to be most resolvable, and the study's own framing rule forbids it.
Is 54.8% good or bad? It is the first receipt-bound per-engine number to compare anything against. Judged against what a citation implies — that the link proves the claim — it is low.
Can a brand raise its own supported rate? Partly. Being fetchable at collection time and stating claims in checkable form are publisher-side. Answer format and model extrapolation are not.
Where do the numbers come from?
The live aggregate at machinerelations.ai/measurement/answer-source-fidelity, read September 19, 2026, alongside the current Index release mri_score_v2.0+2026-09-19+0cad03121f60.