Gemini Cited No Source in 51.6% of Its Returned Answer Runs
The Machine Relations Index now publishes per-engine run denominators. Perplexity cited a source in 99.97% of returned runs. Gemini cited one in 48.4%.
A brand checking its AI visibility one engine at a time is reading six instruments that disagree about whether an answer cites anything at all. Until today that disagreement could not be measured from public data. The Machine Relations Index release dated September 28, 2026 is the first to publish a run-level denominator beside the run-level count for every engine it tracks, and the spread it exposes is 51.5 percentage points wide.
Perplexity returned 3,846 answer runs in the window and 3,845 of them cited at least one source. Gemini returned 3,838 and 1,859 of them cited at least one source. Same prompt basket, same 135 days, same index. One engine effectively always attributes; the other attributed in fewer than half its returned runs.
That is not a ranking of how good the engines are, and the Index deliberately declines to grade it. It is the denominator every per-engine visibility number is quietly divided by, and it has been invisible.
The six engines, in full
Read from the release manifest at mri-release-manifest.json, under engine_roster. Release mri_score_v2.0+2026-09-28+054b584266b8, window May 10 through September 28, 2026, 135 observed days, 1,010 monitored prompts, 17,868 answer runs, 139,633 citation events, 24,639 distinct cited domains.
| Engine | Returned runs | Runs that cited a source | Share | Runs that cited nothing |
|---|---|---|---|---|
| Perplexity | 3,846 | 3,845 | 99.97% | 1 |
| ChatGPT | 3,745 | 3,184 | 85.0% | 561 |
| Claude | 3,838 | 3,190 | 83.1% | 648 |
| Google AI Mode | 1,532 | 1,236 | 80.7% | 296 |
| Google AI Overviews | 1,069 | 653 | 61.1% | 416 |
| Gemini | 3,838 | 1,859 | 48.4% | 1,979 |
| All six | 17,868 | 13,967 | 78.2% | 3,901 |
Across the whole release, 3,901 returned answer runs — 21.8% of them — produced no citation of any kind. More than half of those belong to one engine.
What the denominator actually counts
The field names are runs_total and runs_observed, and their definitions matter more than the headline.
runs_total counts answer runs that the engine returned successfully, against the same eligible prompt set and the same window bounds as every other figure in the release. A run that failed, timed out or was never scheduled is not in it. runs_observed counts how many of those returned runs have at least one row in the Index's citation ledger — not a self-reported flag on the run, but the same ledger that every other observation figure in the release is drawn from.
So the difference between the two columns is a specific, narrow event: the engine came back with an answer, and that answer carried no citation the Index could record.
This is exactly the distinction we argued for in September: a successful answer with zero citations is a measured result and belongs in the rate; a timeout or an unscheduled observation is a coverage problem and belongs somewhere else. The two had to be separated before a number like 48.4% could mean anything. They now are, in public, per engine.
Why the day count could not have told you this
The Index has published per-engine day counts for some time, and a study of them last week showed how far apart the engines' observation windows sit. In today's release Perplexity was observed on 135 of 135 days, Claude on 134, ChatGPT on 131, Gemini on 113, Google AI Mode on 93 and Google AI Overviews on 76.
Those are real and they are a different fact. A day counts once whether the engine cited on all of that day's queries or on one of them, so a nearly empty day and a full one read identically in the day column. The day figure also divides by a window that includes days an engine never ran at all, which means a single number was carrying two unrelated shortfalls and resolving neither.
Gemini is the clean illustration. Its day figure reads 113 of 135, or 83.7%. Its run figure reads 1,859 of 3,838, or 48.4%. Both are true, both are about the same engine in the same window, and they are thirty-five points apart. A reader who has only the day column will conclude Gemini is broadly comparable to ChatGPT and Claude. On the run column it is not close to either.
Neither column replaces the other. A day count survives an engine going dark and recovering, which a run count dilutes; a run count catches a thin day, which a day count cannot see. The release now carries both.
Why Gemini's number has the shape it does
The mechanism is documented and it is not a fault.
Grounding a Gemini answer in web search is an opt-in tool rather than default behaviour. Google's own API documentation describes Search grounding as a capability a caller enables per request, after which the response carries grounding metadata and source links (Google AI for Developers); the enterprise documentation describes the same arrangement, where grounding to Google Search or to private data is configured rather than assumed (Vertex AI). A model asked a question it can answer from its own parameters will do so, and a parametric answer has nothing to cite. Gemini is described by its maker as a general reasoning model first (Google DeepMind), not as a search product.
The same structure applies to the other API engines with different defaults. OpenAI exposes web search as a tool whose result set can be empty for a given call (OpenAI). Anthropic's web search tool is likewise a capability a request opts into (Anthropic docs), introduced as an addition to the model rather than a replacement for it (Anthropic). Perplexity's product is retrieval first and its API is documented around search-backed answers (Perplexity), which is the most plausible reading of a 99.97% rate.
The underlying architecture makes the separation explicit. Retrieval-augmented generation pairs a retriever with a generator precisely so the two can be varied independently (Lewis et al., 2020). When the retriever contributes nothing, the generator still answers. An uncited answer is that architecture working as designed, not evidence that the engine failed.
The two Google surfaces are a third case again
Google AI Mode and AI Overviews are not API calls, and their uncited runs mean something different a third time.
AI Overviews appear when Google's systems determine generative AI is especially helpful for a query, and the feature is available in a named set of countries and languages rather than universally (Google Search Help). Google has described that conditional behaviour since the feature's general rollout (The Keyword), and its guidance to site owners treats AI features as an appearance that may or may not be present for any given search (Google Search Central).
A collection run against a results page that loaded correctly and simply had no AI surface on it is a successful run with nothing to cite. It lands in the uncited column beside a Gemini answer that chose not to retrieve, and the two have almost nothing in common.
The Index refuses to grade these counts, and that is the point
The release publishes runs_total and runs_observed as raw counts and attaches no status, no threshold and no health signal to the ratio between them. The reason is stated in the instrument's own build logic: an answer that cites nothing is a real observation on some engines and a blind lane on others, and nothing available at this layer can tell those two apart.
That restraint is what makes the table above usable. A ratio that had been graded would be asserting a judgment the data cannot support — that Gemini's 1,979 uncited runs are 1,979 deliberate parametric answers, or that Google AI Overviews' 416 are 416 absent surfaces. Some are; nobody can currently say how many. Published ungraded, the counts do the one job they can do honestly, which is to tell a reader what this release is comparable across.
Measurement disciplines that have been at this longer draw the same line. Statistical agencies treat an item with no reported value as missing data requiring separate handling, not as a reported value of zero (U.S. Census Bureau), and the standard engineering treatment of missing observations is to characterise the mechanism before any rate built on them is trusted (NIST/SEMATECH).
What a search team should do with a per-engine zero
Three things change once the denominator is visible.
Stop reading a low Gemini number as a verdict on your authority. If fewer than half of an engine's returned answers cite anything, then roughly half its runs could never have contained your domain regardless of how well your pages are written, linked or structured. The honest statement about a brand absent from Gemini is that it was absent from the 1,859 runs that cited anything, and that the other 1,979 carry no information about it either way.
Normalise before comparing engines, and say which denominator you used. A raw per-engine citation count divided by returned runs and the same count divided by runs that cited anything are different metrics, and on this release they differ by a factor of two on Gemini and not at all on Perplexity. Any dashboard that presents one cross-engine visibility percentage is blending them.
Treat uncited runs as missing information rather than as negative evidence. An engine that cites in 48.4% of returned runs is a low-information instrument for absence claims, whatever its volume. Confidence that a brand is genuinely not being cited should scale with the share of runs that cited anything, not with the number of runs collected.
None of this makes Gemini a worse engine or Perplexity a better one for a reader. It makes them incomparable on the axis most visibility tools quietly treat as shared.
What this does not say
It does not say what proportion of uncited runs are deliberate parametric answers versus collection artefacts. That split is not available in the published data for any engine, and the Index declines to estimate it.
It does not measure whether the cited sources were accurate or relevant, which is a separate instrument with separately published per-engine rates.
It does not generalise beyond the surfaces measured. Four of the six are API surfaces on named model versions — ChatGPT on gpt-4.1-mini-2025-04-14, Claude on claude-haiku-4-5-20251001, Gemini on gemini-3.8-flash, Perplexity on sonar — and the consumer applications carrying those brand names may behave differently, with different retrieval defaults. The two Google surfaces are the public products themselves.
And it is one window. Retrieval defaults are configuration, and configuration changes.
FAQ
Does a 48.4% rate mean Gemini is citing sources less often than it used to? This release does not answer that. It publishes the rate for one 135-day window, not a trend. The run-level pair first appears in the September 28, 2026 release, so no earlier release carries a comparable figure.
Is an uncited answer the same as a failed answer?
No, and the distinction is built into the denominator. runs_total counts only runs the engine returned successfully. A failure, a timeout or an unscheduled observation is excluded before this rate is computed.
Which engine should a brand optimise for on this evidence? None, on this evidence alone. The table describes how often each engine attributes at all. It says nothing about which engine reaches a given audience, and attribution frequency is not the same as commercial value.
Why is Perplexity at 99.97% rather than 100%? One returned run out of 3,846 carried no recordable citation. The Index publishes the count rather than rounding it away.
Where do these numbers come from?
The engine_roster block of the public release manifest at https://machinerelations.ai/data/mri-release-manifest.json, release mri_score_v2.0+2026-09-28+054b584266b8. Every figure in this piece is derived from runs_total, runs_observed and days_observed in that block, and the shares are computed here rather than published there.