Gemini Citation Change Log: 452 of the Top 500 Domains, Recorded Six Sources Deep
A dated record of what Gemini cites and what changes. Entry one: Gemini reaches 452 of the 500 most-cited domains, the second-widest head overlap of six engines behind only Perplexity, and cites the leading source in 86 of 93 published segments, while AuthorityTech's own daily panel was recording only 6 of its roughly 10.8 sources per answer.
This log is a dated record of what Gemini cites and what changes, read from the Machine Relations Index and AuthorityTech's own daily panel. Entry one: Gemini reaches 452 of the 500 most-cited domains in AI answers — second only to Perplexity's 490 — and cites the single most-cited source in 86 of the 93 published question segments, second only to Perplexity's 93. The panel that tracks it day to day was recording a mean of 6 sources per answer against the roughly 10.8 the engine actually returns, disclosed here before any churn number built on it.
Every entry states the measurement, the release or window it came from, the primary source, and what a publisher can check. Entries are appended; nothing is rewritten. Entry one draws on the Machine Relations Index, which observes six engines — Perplexity, ChatGPT, Claude, Gemini, Google AI Mode and Google AI Overviews — against a fixed basket of buying and research questions and publishes a machine-readable release daily. Entry one also covers a methodological finding about AuthorityTech's own daily panel that has to be on the record before this log uses that panel for anything: the panel's citation log keeps only the first six sources per answer, and after Perplexity, Gemini is the engine it caught worst.
Entry 1 — September 22, 2026: the baseline
Release mri_score_v2.0, contract machine_relations_index_public_view_v2.0, generated September 22, 2026. Window May 10 to September 22, 129 days, 16,307 observed answer runs, 34,720 cited-domain observations, 22,828 cited domains. A subject category paired with a question type forms a stratum, and a stratum publishes a rate only after clearing an evidence floor of 10 observed runs across 7 distinct dates; 93 of the 151 measurable strata are published today.
The Index records, for each domain, which engines cited it at least once in the window. It does not publish a per-engine citation rate, so everything below is presence and counts, not per-engine rates — the same boundary this log's companion ChatGPT and Perplexity entries run under.
Second-widest head overlap of six engines
| Engine | Top 25 | Top 100 | Top 250 | Top 500 |
|---|---|---|---|---|
| Perplexity | 25 | 96 | 244 | 490 |
| Claude | 23 | 94 | 237 | 463 |
| Gemini | 23 | 89 | 227 | 452 |
| Google AI Mode | 24 | 91 | 227 | 444 |
| Google AI Overviews | 24 | 89 | 203 | 378 |
| ChatGPT | 22 | 88 | 203 | 361 |
Gemini misses 48 of the top 500 most-cited domains in the market — 11 of the top 100, 23 of the top 250. Only Perplexity and Claude reach deeper into the head. Gemini and Google AI Mode are effectively tied at the 250 mark (227 each) and Gemini pulls ahead by the 500 mark (452 against 444), which is the band that separates it from the middle of this field rather than the top.
Second on segment leaders, too
For each of the 93 published strata, take the domain with the highest citation rate in that stratum — the source the engines cite most when buyers ask that category's question — then ask which engines cite that domain anywhere in the window.
| Engine | Cites the leading source, of 93 segments |
|---|---|
| Perplexity | 93 |
| Gemini | 86 |
| Google AI Mode | 81 |
| Google AI Overviews | 79 |
| Claude | 72 |
| ChatGPT | 66 |
Gemini misses the segment leader in 7 of 93 published categories. Combined with the head-overlap table, the pattern across both published composition facts is the same: Gemini is not the widest engine measured, but on every measure that ranks by where the market's evidence actually concentrates, it is the closest follower to Perplexity of the remaining five.
Domain count, stated without an exclusivity claim
Gemini cites 7,118 of the 22,828 domains observed in the window — fewer than Perplexity's 9,986 and Google AI Mode's 7,441, more than Claude, Google AI Overviews and ChatGPT. This log's ChatGPT and Perplexity entries each carry a floor-adjusted exclusive-domain count for their own engine, computed against the specific release each entry used. That figure is not restated here for Gemini: the floor-adjusted count has to be computed against this same September 22 release to be comparable to the numbers above, and that computation has not been done. An unconditioned exclusive-share number would repeat the mistake this log's ChatGPT entry already corrected once — most of an unconditioned count's weight sits on domains cited a single time — so this entry states domain count and stops there rather than publish a number this log would have to correct later.
Community platforms: fourth of six
Twenty-seven classified domains in the Index are community and social platforms — a small set where engines diverge most.
| Engine | Community platforms cited (of 27) |
|---|---|
| Google AI Mode | 20 |
| Perplexity | 18 |
| Gemini | 15 |
| Claude | 10 |
| Google AI Overviews | 10 |
| ChatGPT | 8 |
Gemini sits closer to Google AI Mode and Perplexity's community reach than to Claude, Google AI Overviews or ChatGPT's, echoing where it lands on the two composition tables above.
The one robots.txt token that governs eligibility, and what it does not cover
Unlike ChatGPT, Perplexity or Claude, Gemini has no dedicated fetch-time citation crawler for a publisher to name in robots.txt. Google's common-crawlers documentation, last updated July 14, 2026, lists eleven Google crawler and fetcher tokens; the only one that names Gemini is Google-Extended, which the same page states "doesn't have a separate HTTP request user agent string" — crawling happens under Google's existing user-agent strings, and the robots.txt token exists purely "in a control capacity." A publisher cannot see it arrive in server logs because it never does.
Google's documentation states plainly that Google-Extended governs two things together, not one: whether crawled content may be used to train future Gemini models, and whether it may be used for grounding — "providing content from the Google Search index to the model at prompt time" — in Gemini Apps and in Grounding with Google Search on Vertex AI. Grounding is the mechanism that produces a citation. A publisher-facing explainer of the token puts the practical stakes plainly: "the cost of disallowing it is losing grounding citations in Gemini rather than losing rankings," and the same source notes there is no way in robots.txt to allow grounding while blocking training, or the reverse — one token, one directive, both uses. A second explainer reaches the same reading of Google's own page: there is no Gemini-named crawler to permit, only a training-and-grounding control token that a publisher either allows or disallows in full. Separately, a robots.txt reference for AI crawlers notes that disallowing Google-Extended does not remove a site from Google AI Overviews, which run on the regular Googlebot index rather than this token — a distinction worth stating because the two Google surfaces are easy to conflate and this log tracks Gemini, not AI Overviews.
This is a structural difference from the other engines in this series, not a finding about any specific domain's absence. Where the ChatGPT and Perplexity entries could name individual blocked or permitted domains and check each one's robots.txt, Gemini's single, log-invisible token means that check cannot be run the same way, and this entry does not attempt it.
Grounding also runs on content Google reaches through means other than crawling. Reuters reported in February 2024 that Google struck a roughly $60-million-a-year licensing agreement giving it access to Reddit's Data API, separate from and in addition to whatever Google-Extended permits. Community-platform reach in the table above is not explained by crawler permission alone.
What the daily panel was actually recording
AuthorityTech runs its own tracked panel alongside the Index: 35 fixed queries, five retrieval surfaces, once a day, the same instrument behind this log's ChatGPT and Perplexity entries and the site's 89-day per-engine churn measurement. That page's September 22 correction found the daily log stores only the first six sources returned by each engine, and the cap does not bite every engine equally. After Perplexity, it bites Gemini hardest of the five.
Measured against the stored per-run answer artifacts, which preserve each provider's complete returned source list — available from September 1, 2026 onward, not for the full panel history — Gemini returns a mean of 10.8 source URLs resolving to 8.8 distinct domains per answer. The daily log recorded a maximum of six. Gemini's answers exceeded six sources on 175 of 193 runs in that window — 90.7% truncated — against Perplexity's 350 of 350 (100%), Claude's 16%, ChatGPT's 12%, and Google AI Mode's 0%.
The practical consequence, recomputed rather than assumed small: day-over-day domain churn on Gemini's complete source lists for the September 1–22 window is 42.8%, against 46.0% on the panel's six-source cap — an overstatement of 3.2 percentage points, in the direction the truncation predicts, with the cross-engine churn ranking unchanged. That correction and its full recomputation table live on the churn page linked above; this log will not restate the panel's older 89-day retention, drop-out, and core-domain figures as if they described the complete list, because the complete-list artifacts do not reach back past September 1 and the churn page says so explicitly. Any figure in this log drawn from the panel is labelled by which list produced it.
Where other measurements land, and why they stay off this log
Third-party trackers publish their own Gemini citation-depth numbers, on their own prompt baskets, and they do not agree with each other or with this log's 10.8. Profound reports 6.6 sources per answer against Google AI Mode's 15.2. GetMint's three-engine study of 1,152 answers reports Gemini citing 46 instances and 10 unique domains per answer, driven heavily by Reddit and YouTube. SoRank's citation study puts Gemini at 11.1 sources per answer, the highest of five models it tracked, rising from 4.1 to 7.2 over its measurement period. A 1,250-prompt Indonesian-market study found Gemini's Search grounding fires on only 34.2% of consumer prompts at all, and cites 7.2 sources per answer when it does. A fifth source states a 4-to-8 range and separately confirms the Google-Extended grounding dependency above.
Six figures — 6.6, 7.2, 10.8, 10, 11.1, and a 4-to-8 range — from six instruments measuring six different prompt baskets over six different windows. None of them is wrong on its own terms and none of them is comparable to this log's 10.8 without knowing the query basket behind it. This log's number comes from a fixed panel of 35 buying and research queries measured daily since September 1, checked against the complete per-run artifact rather than a display-layer or third-party scrape; it is stated as what it is, not as an industry consensus figure, because no such consensus exists in the public data this entry could locate.
How to read this log
Each future entry compares a release or a panel window against the one before it and reports what moved. Four numbers are the ones to watch:
- Head overlap. 227 of the top 250, 452 of the top 500.
- Segment leaders cited. 86 of 93.
- Community platforms. 15 of 27.
- Panel-recorded source depth. 8.8 distinct domains per answer on the complete list, against a six-item cap in the daily log; watch for the cap narrowing as the answer-evidence store's history lengthens past September 1.
A change in any of them is an entry. A release with none is recorded as no change.
What a publisher should take from entry one
Gemini is the clearest second-place engine measured so far in this series. It does not lead any of the three composition measures this log tracks — head overlap, segment leaders, community reach — but it is closer to Perplexity's numbers on each of them than any of the other four engines, and it pulls ahead of Google AI Mode in the top-500 band despite matching it almost exactly at the top-250 band. A brand chasing broad AI-search presence should treat Gemini as the second target after Perplexity, not an afterthought behind ChatGPT.
The measurement finding cuts the other way, against AuthorityTech's own panel rather than the engine. A dashboard, a competitor tracker, or an internal panel that caps a citation array at six entries per answer will report close to the true count for Google AI Mode, understate ChatGPT and Claude modestly, and understate Gemini by nearly half — 8.8 true domains against a 6-item cap, truncated on 9 of every 10 runs. AuthorityTech's own panel did this for the whole window through September 1, and only caught it by checking the complete artifact against the six-item log.
Watch list for entry 2
- Whether Gemini's head-overlap or segment-leader numbers move at the next Index release — the release updates daily, so entry 2 will report on genuine movement, not a fixed weekly cadence.
- Whether a floor-adjusted exclusive-domain count for Gemini gets computed against a specific release, making that comparison publishable without the caveat in entry one.
- Whether the complete per-run artifact history extends further back than September 1, letting the panel restate churn over a longer window than 22 days.
- Whether any additional top-500 domain moves into or out of Gemini's cited set.
Frequently asked questions
Is Gemini the second-strongest engine on every measure in this series? Not on every measure — it is second on head overlap and segment leaders, fourth of six on community-platform reach, and cites fewer distinct domains overall than Perplexity or Google AI Mode. The consistent pattern is that wherever this log ranks engines by where the market's citation evidence concentrates, Gemini is the nearest follower to Perplexity among the remaining five.
Why does this entry not publish an exclusive-domain-share number for Gemini the way the ChatGPT and Perplexity entries did for their own engines? Because that number has to be computed against the same release as the rest of this entry's figures to be comparable, and that computation was not available when this entry was written. Publishing an unconditioned exclusive-share figure now would repeat the error this log's ChatGPT entry already corrected once — the raw count is dominated by domains cited only once — so this entry states total domain count and leaves the floor-adjusted figure for a later entry.
Is the six-source cap a Gemini problem or a measurement problem? A measurement problem, and specifically AuthorityTech's own daily panel's problem, not the Machine Relations Index's — the Index used for entry one's composition figures is not affected by this cap. It is disclosed here, before this log uses the daily panel for anything, so that no churn or stability number in a future entry needs a correction after the fact.