ChatGPT Citation Change Log: 55% of What It Cites, No Other Engine Cites
A dated record of what ChatGPT cites and what changes, read from the Machine Relations Index. Entry one: ChatGPT has the weakest overlap with the index head of any of the six engines and the largest share of citations no other engine makes.
ChatGPT cites 4,877 of the 22,179 domains observed in AI answers, which makes it the fifth-widest of six engines. That number on its own is misleading. More than half of those domains — 2,708 of 4,877, or 55.5% — are cited by no other engine in the same window. No other engine is close: Claude's exclusive share is 27.9%, Google AI Overviews' is 16.0%.
ChatGPT does not cite less than the other engines. It cites somewhere else.
This page is a dated record of what ChatGPT cites and what moves. Every entry states the measurement, the release it came from, the primary source for any policy it references, and what a publisher can check. Entries are appended; nothing is rewritten. The measurements come from the Machine Relations Index, which observes six engines — ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity — against a fixed basket of buying and research questions, and publishes a machine-readable release daily.
Entry 1 — September 18, 2026: the baseline
Release mri_score_v2.0, contract machine_relations_index_public_view_v2.0, generated September 18, 2026. Window May 10 to September 18, 125 days, 15,782 observed answer runs, 124,397 citation events, 22,179 cited domains. A subject category paired with a question type forms a stratum, and a stratum publishes a rate only after clearing an evidence floor of 10 observed runs across 7 distinct dates; 85 of the 151 measurable strata are published today.
The Index records, for each domain, which engines cited it at least once in the window. It does not publish a per-engine citation rate, so everything below is presence and counts, not per-engine rates.
ChatGPT's citation set is narrow at the top and wide at the tail
| Engine | Domains cited | Cited by that engine only | Exclusive share |
|---|---|---|---|
| Perplexity | 9,497 | 4,246 | 44.7% |
| Google AI Mode | 7,441 | 3,382 | 45.5% |
| Gemini | 7,079 | 2,837 | 40.1% |
| Claude | 5,138 | 1,431 | 27.9% |
| ChatGPT | 4,877 | 2,708 | 55.5% |
| Google AI Overviews | 2,279 | 365 | 16.0% |
Now hold that against the head of the index — the most-cited domains, the ones every visibility programme targets.
| Engine | Top 25 | Top 100 | Top 250 | Top 500 |
|---|---|---|---|---|
| Perplexity | 25 | 96 | 243 | 489 |
| Claude | 23 | 94 | 236 | 462 |
| Gemini | 23 | 89 | 227 | 455 |
| Google AI Mode | 24 | 93 | 227 | 448 |
| Google AI Overviews | 24 | 90 | 205 | 385 |
| ChatGPT | 22 | 88 | 203 | 356 |
ChatGPT has the lowest overlap with the index head of any of the six engines, at every band. Google AI Overviews cites half as many domains in total and still reaches more of the top 500.
That combination is the finding. An engine with a small, head-concentrated citation set is a narrow engine. An engine with a large, head-concentrated set is a broad one. ChatGPT is neither: it skips 144 of the 500 most-cited domains in the market and spends its breadth on 2,708 domains nobody else cites at all.
The community gap
The Index classifies cited domains by the role they play as sources. Only 3,556 of the 22,179 carry a classification; the rest are uncategorised, so shares within the classified set are the honest comparison, not shares of everything.
Twenty-seven classified domains are community and social platforms. That small set is where the engines diverge most sharply.
| Engine | Community platforms cited (of 27) | Editorial share of its classified set |
|---|---|---|
| Google AI Mode | 20 | 30.0% |
| Perplexity | 18 | 32.5% |
| Gemini | 15 | 31.4% |
| Claude | 10 | 35.5% |
| Google AI Overviews | 10 | 31.9% |
| ChatGPT | 8 | 37.9% |
ChatGPT cites the fewest community platforms of any engine and carries the highest editorial-publication share of its classified set. Whatever else is true of it, this engine leans to published media and away from forums harder than the other five.
Where that shows up: the leading source in 24 of 85 segments
For each of the 85 published strata, take the domain with the highest citation rate in that stratum — the source the engines cite most when buyers ask that category's question. Then ask which engines cite that domain anywhere in the window.
| Engine | Cites the leading source, of 85 segments |
|---|---|
| Perplexity | 85 |
| Gemini | 78 |
| Google AI Mode | 76 |
| Google AI Overviews | 74 |
| Claude | 64 |
| ChatGPT | 61 |
Perplexity cites the leading source in every published segment. ChatGPT misses 24 of them — and 21 of those 24 are the same domain: reddit.com, leading segments including is_x_worth in AI infrastructure (32.1%), consumer finance (31.8%) and consumer products (27.9%), and best_x in education and training (34.0%).
The other three are teramind.co on AI security and privacy how_choose (26.2%), unmannedsystemstechnology.com on deep-tech hardware top_list (28.3%), and erpfocus.com on enterprise software x_vs_y (20.2%).
The 47 head domains absent from ChatGPT, and what their robots files say
Forty-seven of the 250 most-cited domains do not appear in a single ChatGPT run in this window. Reading each publisher's own robots.txt on the same day is a cheap check: it establishes the publisher's stated policy. It never establishes cause, and several of these domains contradict it.
| Rank | Domain | Class | Rate | robots.txt, read 2026-09-18 |
|---|---|---|---|---|
| 1 | reddit.com | Community | 12.75% | User-agent: * / Disallow: / — everything blocked |
| 18 | substack.com | Community | 1.58% | No OpenAI agent named; no blanket block |
| 23 | facebook.com | Community | 1.29% | GPTBot restricted to path rules, not a full block |
| 42 | nytimes.com | Editorial | 0.89% | GPTBot, ChatGPT-User and OAI-SearchBot each Disallow: / |
| 54 | pcmag.com | Editorial | 0.74% | All three OpenAI agents in a blocked group |
| 107 | cnet.com | Editorial | 0.47% | All three OpenAI agents in a blocked group |
| 108 | wsj.com | Editorial | 0.47% | No OpenAI agent rule |
| 174 | investopedia.com | Unclassified | 0.36% | OpenAI agents allowed but for /thmb/ |
The New York Times blocks all three OpenAI crawlers by name and is absent, which is consistent — and it is also in active litigation with OpenAI and Microsoft, with filings through early September 2026. CNET and PCMag name the same three agents in blocked groups. OpenAI's crawler documentation distinguishes GPTBot (training), OAI-SearchBot (search index) and ChatGPT-User (user-triggered fetch), so a publisher can block one and permit another; these three block all of them.
Then the contradictions. Investopedia permits every OpenAI agent except an image path and is absent anyway. The Wall Street Journal names no OpenAI agent at all — its parent has a reported $250 million five-year agreement with OpenAI, the largest publisher deal on record — and is absent anyway. Substack and Facebook impose no relevant block and are absent.
Permission is necessary. It is plainly not sufficient.
The Reddit inversion
Reddit is the most-cited domain in AI answers at 12.75%, cited in 2,012 of 15,782 runs. Its robots.txt is eight lines and blocks every crawler: User-agent: *, Disallow: /, pointing at Reddit's Public Content Policy for access terms.
So the four engines that cite Reddit are not reaching it by crawling. Google reached it by contract — a licensing agreement reported at roughly $60 million a year — and OpenAI signed a comparable agreement, reported at around $70 million a year, in May 2024. Reddit has since sued Anthropic and pursued other platforms over unlicensed use.
The engine holding the licence is the one that does not cite the source. That is the inversion this log will keep watching.
Where the measurements disagree, and that stays on the record
Third-party panels do record ChatGPT citing Reddit. Treg's citation panel for the query "what sources does chatgpt cite", sampled September 18, 2026, returns 429 cited domains with reddit.com second at 204 mentions. And ZeroClick Labs reported in August 2026 that ChatGPT's visible Reddit citations fell roughly 86% between August 8 and August 17 while retrieval held steady — a display-layer change, on their reading, rather than a block.
Those are different instruments measuring different things over different prompt baskets, and the disagreement is the honest state of the record. The claim that survives all of them is narrow and worth stating precisely: across this basket of buying and research questions over 125 days, the Index observed no ChatGPT run citing reddit.com. Whether that is access, ranking or display is not settled by any public measurement, and this log will not pretend otherwise.
How to read this log
Each future entry compares a release against the one before it and reports what moved. The figures above are the fixed baseline for the September 18, 2026 release; the next entry reads the following weekly release against them. Four numbers are the ones to watch:
- Head overlap. 203 of the top 250, 356 of the top 500.
- Exclusive share. 55.5% of ChatGPT's 4,877 domains.
- Community platforms. 8 of 27.
- Segment leaders cited. 61 of 85.
A change in any of them is an entry. A release with none is recorded as no change, because a log that only reports movement teaches readers to mistake quiet periods for missing data.
What a publisher should take from entry one
Being in the index head is not the same as being in ChatGPT. One hundred and forty-four of the 500 most-cited domains in AI answers do not appear in a ChatGPT run in this window, and the reasons are not uniform: some blocked the crawlers, some signed agreements and are absent anyway, some did neither.
The practical reading is that per-engine presence is a measurement, not an inference. A brand tracking a single aggregate "AI visibility" number cannot see any of the above, because the engines disagree with each other more than they agree — and the engine most people mean when they say AI search is the one that agrees with the others least.
Watch list for the next entry
- Whether ChatGPT's head overlap moves in either direction at the next release.
- Whether reddit.com appears in any ChatGPT run, and whether the OpenAI agreement is renewed, expanded or lapses.
- Whether the Wall Street Journal and Investopedia — permitted and absent — change state.
- Whether the New York Times litigation produces a ruling that changes either party's access posture.
- Whether TechRadar, absent from both Google Search surfaces with no relevant robots rule, resolves or persists.
Frequently asked questions
Does this mean ChatGPT cites fewer sources than other engines? No. It cites 4,877 distinct domains, more than Google AI Overviews' 2,279. The distinctive thing is the composition: the lowest overlap with the most-cited domains in the market, and the highest share of citations no other engine makes.
Is a blocked robots.txt the reason a domain is absent from ChatGPT?
Not reliably. It matches in several cases — the New York Times, CNET, PCMag — and fails in others, including the Wall Street Journal and Investopedia, both of which permit OpenAI's agents and are still absent. A robots.txt file records a publisher's stated policy, and nothing more.
Why does a third-party tool show ChatGPT citing Reddit when this says it does not? Different prompt baskets, different sampling windows, and a distinction between retrieval and displayed citation. Both readings are on this page, unreconciled, because no public measurement settles it.
What does "leading source in a segment" mean? A segment is one subject category paired with one question type. Its leading source is the domain with the highest citation rate in that segment across all six engines. ChatGPT cites that domain, anywhere in the window, in 61 of the 85 published segments.
How often is this updated? On each weekly Index release, and on any dated engine or policy change with a primary source.
Change log
| Date | Entry | Release |
|---|---|---|
| 2026-09-18 | Baseline established: 4,877 domains cited, 55.5% exclusive, 203 of top 250, 8 of 27 community platforms, 61 of 85 segment leaders | mri_score_v2.0, generated 2026-09-18 |
Related: the per-engine source map of the top 25 domains, which covers all six engines for the same release, and the Machine Relations Index Report for September 18, 2026. Reddit's full per-segment record is on its Index domain profile.