Nine in Ten AI Citations Never Name the Source
Across 14,447 readable citations in four AI engines, the cited source's name reached the answer text 8.6% of the time. The rate splits by engine.
A citation and a mention are two different events, and in the Machine Relations Index panel they come apart badly. Across 14,447 citations we can read in full, from four AI answer engines over 139 days, the cited source's own name appeared in the words of the answer 8.6% of the time. The other 91.4% were silent: the engine used the page, linked it in a source layer the reader has to open, and wrote the sentence without ever saying whose page it was.
The rate is not the same on every engine. Gemini named a cited source in 13.2% of its citations. Perplexity named one in 5.9%. That is a 2.2x spread on the same prompts in the same window, which means a tool reporting "citations" as one number across engines is adding up two different outcomes.
What we measured, and on which engines
The source is release mri_score_v2.0+2026-09-26+2b779408cfda of the Machine Relations Index, the neutral public benchmark for which sources AI answer engines cite. The release window runs 2026-05-10 to 2026-09-26 and covers 16,925 observed answer runs against 1,010 monitored buyer prompts across six engines, with 23,978 cited source domains and 132,514 citation events. Counts and methodology are published at the Index.
The measured set is the 96 domains cited in 80 or more of those answer runs — enough evidence for each one to be split four ways by engine. For every pairing of an answer and a measured domain, we recorded two separate outcomes: whether the engine cited that domain, and whether the answer's prose contained that domain's brand name.
"Prose" is doing real work in that sentence. Before matching a name we removed every Markdown link and its anchor text, every URL, and every bare domain string from the answer. An engine that writes its source list into the body of its reply cannot score its own sources as named. That treatment cost ChatGPT 60 of its matches, which is exactly why it is there.
Four engines carry answers we can read this way, and their coverage is high enough to compare:
| Engine | Answer runs in window | Runs with stored answer text | Coverage |
|---|---|---|---|
| Gemini | 3,615 | 3,400 | 94.1% |
| Claude | 3,615 | 3,399 | 94.0% |
| ChatGPT | 3,522 | 3,306 | 93.9% |
| Perplexity | 3,623 | 3,380 | 93.3% |
Nine in ten citations are silent
| Engine | Citations read | Source named in the prose | Rate | Same name in answers that did not cite the domain |
|---|---|---|---|---|
| Gemini | 2,865 | 377 | 13.2% | 1.45% |
| ChatGPT | 2,261 | 238 | 10.5% | 0.34% |
| Claude | 1,820 | 184 | 10.1% | 0.37% |
| Perplexity | 7,501 | 440 | 5.9% | 0.26% |
| Four engines pooled | 14,447 | 1,239 | 8.6% | — |
The right-hand column is the control, and it is the reason the rest of the table means anything. For each engine we counted how often the same brand name turned up in answers from that engine that did not cite the domain. Those rates sit between 0.26% and 1.45%, so a citation raises the odds of the name reaching the prose by a factor of 9 to 39. The association is real. It is also small in absolute terms: on the most generous engine in the set, roughly seven of every eight citations still leave the source unnamed.
Every engine here attaches its sources as a structured layer rather than as words. OpenAI's web search tool returns "answers with sourced citations" in the Responses API. Anthropic's documentation says the response "includes citations for sources drawn from search results" as structured citation blocks. Perplexity describes its answers as web-grounded with built-in citations, and its own citation instruction attaches a marker to each sentence carrying retrieved information. Google describes AI Overviews as a snapshot with links to dig deeper, and documents the surface for site owners under AI features in Search. In all four systems, the sentence and the attribution are produced by different mechanisms, and only one of them is read aloud.
Perplexity's low rate survives the obvious objection
Perplexity cites far more than the others — 2.22 measured domains per answer against 0.68 for ChatGPT and 0.54 for Claude. A wider citation set mechanically dilutes any per-citation naming rate, because an answer names a handful of things however many pages it consulted. So the spread could be arithmetic rather than behaviour.
It is not. Holding the number of measured domains cited in the answer fixed, Perplexity is lowest in every band:
| Measured domains cited in the answer | ChatGPT | Gemini | Claude | Perplexity |
|---|---|---|---|---|
| 1 | 10.5% | 10.2% | 9.9% | 7.8% |
| 2 | 9.2% | 13.6% | 9.2% | 6.5% |
| 3 | 12.7% | 12.1% | 11.8% | 5.3% |
| 4 or more | 11.5% | 16.1% | 15.8% | 5.4% |
The other three engines drift upward as the citation set widens; Perplexity holds flat near 5 to 8%. The independent reading is its attribution style: Perplexity credits with a numbered marker per sentence, so attribution reaches the reader as a footnote rather than as a name in the clause. Gemini goes the other way and pulls further ahead as it cites more, on the longest answers in the panel.
Being named is a property of what you are, not of how good the citation was
The engine split is real but small next to the split by source. When the same measure is run per domain, with each domain's own not-cited control beside it, the range runs from 28.6% to zero on citation counts of 100 or more:
| Domain | Citations read | Named in the prose | Control, not cited | Lift |
|---|---|---|---|---|
| databricks.com | 105 | 28.6% | 0.6% | 44.5x |
| github.com | 142 | 21.8% | 1.1% | 19.5x |
| hubspot.com | 122 | 21.3% | 1.5% | 13.8x |
| crunchbase.com | 122 | 18.0% | 0.2% | 114.7x |
| deloitte.com | 110 | 13.6% | 0.2% | 76.0x |
| gartner.com | 270 | 11.9% | 1.2% | 9.8x |
| zapier.com | 116 | 11.2% | 0.6% | 19.7x |
| reddit.com | 1,672 | 3.5% | 1.0% | 3.6x |
| forbes.com | 559 | 2.1% | 0.1% | 15.4x |
| medium.com | 560 | 4.6% | 3.1% | 1.5x |
| wikipedia.org | 171 | 0.6% | 0.4% | 1.6x |
| youtube.com | 1,126 | 1.2% | 1.1% | 1.1x |
| techtarget.com | 205 | 0.5% | 0.0% | — |
| techradar.com | 454 | 0.7% | 0.0% | — |
| arxiv.org | 405 | 0.0% | 0.0% | — |
| nytimes.com | 132 | 0.0% | 0.0% | — |
| prnewswire.com | 125 | 0.0% | 0.0% | — |
Two patterns fall out, and the second is the one worth budget.
Domains that are the subject of a buying answer get named. Databricks, GitHub, HubSpot and Zapier are products a reader might choose, and when the engine cites their page it is usually recommending the thing, so the name is in the sentence by necessity. Crunchbase, Deloitte and Gartner get named because an answer citing them is often reporting what they said.
Domains that are evidence do not get named, however strong they are. On 405 citations, arxiv.org reached the prose zero times. On 559, Forbes reached it twelve times. TechTarget, TechRadar, The New York Times and PR Newswire are all at or below 0.7%. These are not weak sources — a publication has to be cited hundreds of times to appear in this table at all. They are simply doing a job that does not require the engine to say their name.
Medium, Wikipedia and YouTube are the control working. Their brand words are common enough in ordinary prose that their not-cited baselines are 0.4% to 3.1%, and their lift collapses to between 1.1x and 1.6x. We report them rather than dropping them, because a reader checking this measure should see where it stops discriminating.
Where our figure and the published one disagree
Kevin Indig named this gap the ghost citation, and Semrush measured it with him: 3,981 domain appearances across 115 prompts in 14 countries and four engines, with 61.7% of citations producing no brand mention in the answer, published as the ghost citations study. Their ghost rate is 62%. Ours is 91.4%.
Both point the same direction, and the difference is method, not contradiction. Their unit is a domain appearance on a hand-built prompt set across 14 countries; ours is a domain-and-answer pair on 1,010 continuously sampled US-English buyer prompts, and we strip link anchors, URLs and bare domain strings before matching a name, which removes the easiest way for a name to appear without the engine having written it. A stricter reading of "named" produces a higher ghost rate, which is what a stricter reading should do. Their domain-level finding is ours as well: they report Medium cited and never named, and Wikipedia and Harvard cited but rarely named, which is the evidence-versus-subject split above arriving from a different panel.
The academic work is converging on the same seam from the measurement side. A two-stage framework for citation selection and citation absorption, published April 2026 over 602 controlled prompts and 21,143 search-layer citations with its dataset in the geo-citation-lab collection, reports that Perplexity cites the most sources per prompt while ChatGPT cites fewer with substantially higher influence per fetched page. That is the same shape our conditioned table shows: breadth and depth of citation are separate quantities, and only one of them reaches the reader.
Why Google's two surfaces are not in the comparison
Google AI Mode and Google AI Overviews are in the Index and are excluded from this measure, for two reasons stated rather than hidden.
First, coverage. Stored answer text exists for 63.8% of AI Mode runs and 28.9% of AI Overviews runs in this window, against 93 to 94% for the four engines above. A rate computed on a quarter of a surface's answers describes the quarter we captured.
Second, and decisively, the shape of what is captured. For Google's two surfaces the stored text interleaves source-attribution labels with the answer itself — the site chips Google renders beside a passage arrive as words in the same string. A name can therefore be present without the answer having written it, and no amount of URL stripping separates the two. On the unfiltered measure those two surfaces score 33.0% and 41.8%, which would have topped this table and would have been an artifact of our capture rather than a fact about Google. We publish no naming rate for them until the capture separates chips from prose.
That is the whole reason to state a measure's text treatment out loud. The two highest numbers we computed are the two we will not publish.
What this changes for a brand's team
Stop reading a citation count as a visibility number. It is a supply number: it says your page was in the set the engine assembled the answer from. On this panel, roughly eleven citations buy one appearance of your name in the text a buyer reads, and the exchange rate moves by engine and by what kind of source you are.
Track the two outcomes separately, per engine, because they respond to different work. Being cited responds to being retrievable and extractable on the question. Being named responds to being the entity the category is about — which our own panel work finds is driven by how many independent sources already count you in that category, not by where you place yourself on your own list. The evidence for that sits in the corroboration section of the Index's content-structure study and in our practitioners' write-up of the citation economy; the founder-side argument for why earned coverage moves this and on-page work does not is at why earned media beats content tweaks, and the brand-side measurement framing is at brand mentions and source architecture.
And when a dashboard reports one blended AI visibility figure, ask which of these two events it counted and on which engines. A tool that pools Perplexity's numbered attributions with Gemini's named ones, or that counts a link anchor as a mention, is reporting a mixture whose composition changes every time the engine mix changes.
Methodology and limits
The panel is the Machine Relations Index public view, release mri_score_v2.0+2026-09-26+2b779408cfda, window 2026-05-10 to 2026-09-26. We reproduced the release's own eligibility rules before cutting anything new — public scoring eligible prompts only, the two public panels, the six engines, the release window, and hosted-subdomain resolution at confidence 0.8 or above — and matched the release exactly at 16,925 answer runs and 23,978 cited domains before computing a single naming rate.
- Name matching is lexical. A domain's brand name is taken as its registrable label, matched case-insensitively on word boundaries in the stripped text. It cannot see a brand written differently from its domain, and it over-matches a label that is also an ordinary word. The not-cited control is how you tell which rows suffer from that: Medium at 1.5x lift is unreliable, Databricks at 44.5x is not.
- Association, not cause. Nothing here shows that a citation produces a mention. Both outcomes can follow from the answer being about the thing.
- Coverage bounds each rate. Rates are computed over answers with stored text, 93.3% to 94.1% per engine. Google's two surfaces are excluded on the grounds above.
- Domain level only. These are counts of domains over stated denominators. The Index's public view carries no prompt text, no answer text and no cited URLs, and none of those appear here.
- Evidence floor. The measured set is 96 domains at 80 or more citations each in the window. Thinner domains are not measured, so this describes the head of the cited universe, not its tail.
Paralax is an editorially independent AI search research publication from the team behind the Machine Relations Index, and the Index is the panel used above; our commercial practice is AuthorityTech.
FAQ
Does a citation with no mention still help my brand? It is the mechanism by which the answer gets built, so it is not nothing. But it is not a moment of brand exposure, and on this panel it is not one 91.4% of the time. Price it as supply into the answer, not as a placement a buyer sees.
Which engine is best to be cited by if I want my name in the answer? On this window and these prompts, Gemini at 13.2% and then ChatGPT at 10.5%. Perplexity at 5.9% is the least likely to convert a citation into a name, because it credits by numbered marker.
Why does Perplexity cite so much more? It returns a wider source set per answer — 2.22 measured domains against 0.54 for Claude. The published academic framework on citation selection and absorption reports the same breadth-versus-depth split across a separate panel.
Is 91.4% the same as Semrush's 62% ghost citation figure? Same phenomenon, different measure. We remove link anchors, URLs and bare domain strings before counting a name, and we run a continuously sampled US-English prompt panel rather than a 14-country prompt set. The stricter reading gives the higher ghost rate; both find that most citations are silent.
Can I check any of this? The per-domain citation counts and confidence grades are public at the Index and on each domain's own profile page. The naming rates are ours and are restated here with their denominators so they can be recomputed on a later release.