Which AI Engines Cite Reddit? Four of Six, and the Split Has Held for Two Months
Reddit is the most-cited domain in AI answers and is absent from ChatGPT and Claude. The per-engine map of the top 25 sources, what each engine's robots and licensing layer says, and what changed since July.
Reddit is the most-cited domain in AI answers and two of the six major engines never cite it. In the Machine Relations Index release of September 18, 2026 (release mri_score_v2.0+2026-09-18+8fa38e54dd0a, window May 10 to September 18, 15,782 observed answer runs), reddit.com holds the top citation rate of any domain at 12.75%, cited in 2,012 runs, and appears in the citations of Gemini, Google AI Mode, Google AI Overviews and Perplexity. It appears in zero ChatGPT runs and zero Claude runs. The same four-of-six pattern was measured on July 25 at 11.55% over 9,258 runs. Fifty-five days and 6,500 more runs later, the split has not moved.
That is the headline. The more useful finding is the map around it. Nineteen of the twenty-five most-cited domains are cited by all six engines. Four of the six exceptions line up with a documented robots.txt or licensing decision on one side or the other, and two, TechRadar and Substack, line up with nothing public. This brief reads the whole per-engine map, then checks it against what each platform and each engine has published about access.
The per-engine map of the top 25
The Index records, for every domain, which engines cited it at least once in the window. It does not publish a per-engine rate, so the table below says presence and nothing more. Engines: C = ChatGPT, A = Claude, G = Gemini, M = Google AI Mode, O = Google AI Overviews, P = Perplexity.
| Rank | Domain | Source class | Citation rate | Runs cited | Grade | Engines present |
|---|---|---|---|---|---|---|
| 1 | reddit.com | Community | 12.75% | 2,012 | A | G M O P |
| 2 | youtube.com | Platform | 9.07% | 1,431 | A | all six |
| 3 | linkedin.com | Community | 6.81% | 1,075 | A | all six |
| 4 | medium.com | Editorial | 5.11% | 807 | A | all six |
| 5 | forbes.com | Editorial | 4.11% | 649 | A | all six |
| 6 | nih.gov | Academic/government | 3.07% | 485 | A | all six |
| 7 | gartner.com | Analyst | 2.88% | 454 | A | all six |
| 8 | techradar.com | Editorial | 2.84% | 448 | A | C A G P |
| 9 | g2.com | Market database | 2.80% | 442 | A | all six |
| 10 | arxiv.org | Academic/government | 2.60% | 410 | A | all six |
| 11 | yahoo.com | Editorial | 1.94% | 306 | A | C A M O P |
| 12 | microsoft.com | Vendor | 1.93% | 305 | A | all six |
| 13 | ibm.com | Vendor | 1.93% | 304 | A | all six |
| 14 | landbase.com | Vendor | 1.91% | 301 | A | all six |
| 15 | healthline.com | Unclassified | 1.88% | 297 | B | all six |
| 16 | crunchbase.com | Market database | 1.72% | 271 | B | all six |
| 17 | nerdwallet.com | Editorial | 1.60% | 253 | B | C A M O P |
| 18 | substack.com | Community | 1.58% | 249 | B | A G M O P |
| 19 | techtarget.com | Editorial | 1.50% | 237 | B | all six |
| 20 | amazon.com | Editorial | 1.50% | 236 | B | all six |
| 21 | paloaltonetworks.com | Vendor | 1.37% | 216 | B | all six |
| 22 | dev.to | Community | 1.32% | 209 | B | all six |
| 23 | facebook.com | Community | 1.29% | 203 | B | G M O P |
| 24 | sentinelone.com | Vendor | 1.27% | 200 | B | all six |
| 25 | deloitte.com | Analyst | 1.24% | 196 | B | all six |
Six exceptions, three shapes. Reddit and Facebook are missing from ChatGPT and Claude. TechRadar is missing from both Google Search surfaces and nowhere else. Yahoo and NerdWallet are missing from Gemini only, and Substack is missing from ChatGPT only. Perplexity is present on all twenty-five.
How wide each engine's source layer is
The gap widens below the head. Of the top 100 domains, 69 are cited by all six engines. Of the top 250, 145. Of the top 1,000, 306. Across the whole universe of 22,179 cited domains, 347 are cited by all six, and 14,969 (about two thirds) were cited by exactly one engine in the window.
Counted per engine, the number of distinct domains each one cited at least once in the window:
| Engine | Distinct domains cited | Domains only this engine cited | Present among the top 1,000 |
|---|---|---|---|
| Perplexity | 9,497 | 4,246 | 967 |
| Google AI Mode | 7,441 | 3,382 | 826 |
| Gemini | 7,079 | 2,837 | 871 |
| Claude | 5,138 | 1,431 | 881 |
| ChatGPT | 4,877 | 2,708 | 606 |
| Google AI Overviews | 2,279 | 365 | 638 |
Two readings follow. Perplexity cites the widest set of sources by every measure and is absent from only four of the top 100 (itpro.com, tracxn.com, researchgate.net, 6sense.com). ChatGPT and Google AI Overviews cite the narrowest slices of the head: ChatGPT is absent from 394 of the top 1,000 and AI Overviews from 362, against 33 for Perplexity. Claude sits in between in an unusual way. It cites fewer distinct domains than Gemini or AI Mode, but it covers the head almost as thoroughly as Perplexity, present on 881 of the top 1,000.
By source class among the top 500, the same two engines are the narrow ones. ChatGPT is absent from 26 of the 97 editorial publications, 16 of the 48 market databases and 21 of the 101 vendor-owned domains in that band. AI Overviews is absent from 25 of the editorial publications and 16 of the market databases. Gemini is absent from none of the 101 vendor-owned domains, and Perplexity from three of the 97 editorial publications.
Where ChatGPT and Claude both go dark
Nine domains in the top 250 are cited by neither ChatGPT nor Claude: reddit.com (1), facebook.com (23), nytimes.com (42), pcmag.com (54), instagram.com (56), cnet.com (107), wsj.com (108), everydayhealth.com (152) and investopedia.com (175). Their robots.txt files, read on September 18, 2026, explain most of that list without any inference about ranking.
Reddit's robots.txt disallows every user agent from every path and points to its Public Content Policy. Reddit announced that change in June 2024 as an enforcement step: crawlers without an agreement should not be accessing Reddit data, and, as its chief legal officer told The Verge, "allow" in robots.txt never meant free use. Google has had access to Reddit's Data API since February 2024, which is the licensing layer on Google's side. OpenAI announced its own Reddit Data API partnership in May 2024, to "bring enhanced Reddit content to ChatGPT." In this panel, that agreement is not visible as a single ChatGPT citation. OpenAI's own crawler documentation says why that can be true at the same time: sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers," and Reddit's file opts out every bot by name. Reddit sued Anthropic in June 2025, alleging that Anthropic accessed Reddit data without a licence. Claude's web search was reported in March 2025 to run on Brave's index, and Reddit's July 2024 robots change cut off search engines without an agreement: Microsoft confirmed at the time that Bing stopped crawling Reddit after the July 1 update.
Facebook's robots.txt is the opposite shape and lands in the same place. It names GPTBot, ClaudeBot, PerplexityBot and Google-Extended with path rules, and disallows everything for any agent it does not name. OAI-SearchBot and Claude-SearchBot, the two agents that decide search-answer inclusion for ChatGPT and Claude, are not named, so they fall under the wildcard block. Facebook is present on the three Google surfaces and Perplexity and absent from the other two.
The New York Times disallows GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, PerplexityBot and Perplexity-User outright while allowing Googlebot with path rules; it sued OpenAI and Microsoft in December 2023. In the Index it is present on Google AI Mode, Google AI Overviews and Perplexity, and absent from ChatGPT, Claude and Gemini. PCMag and cnet.com carry the same full OpenAI, Anthropic and Perplexity blocks; PCMag's parent Ziff Davis sued OpenAI in April 2025. Instagram blocks GPTBot, ClaudeBot, PerplexityBot and Google-Extended and disallows all other agents. Everydayhealth.com blocks every OpenAI, Anthropic and Perplexity agent by name.
Seven of the nine are explained by the file alone. The two that are not are the interesting ones. wsj.com blocks all agents by default but allows GPTBot, OAI-SearchBot and ChatGPT-User by name, consistent with the News Corp agreement that gives OpenAI permission to display News Corp content in answers. Investopedia allows the three OpenAI agents and blocks Anthropic's. Both are absent from ChatGPT in this panel anyway. A permission is a necessary condition for a citation, not a sufficient one.
The file is not a law in the other direction either. CNBC disallows GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and Claude-User by name and is cited by both ChatGPT and Claude, and quora.com blocks every OpenAI and Anthropic agent by name and is cited by Claude. User-triggered fetchers such as ChatGPT-User and Perplexity-User are documented by their vendors as not always bound by robots.txt, and answers can cite a page the engine reached through a search index, a licensed feed or a cached copy rather than a fresh crawl. Presence in the Index is a measurement of citations, not of crawling.
Gemini is not Google Search
Google runs three of the six engines, and they do not cite the same sources. Within the top 250, fourteen domains are cited by both Google AI Mode and Google AI Overviews and never by Gemini: yahoo.com, nerdwallet.com, techcrunch.com, nytimes.com, sciencedirect.com, cnbc.com, instagram.com, tech-insider.org, projectstartups.com, usnews.com, zoominfo.com, stackmatix.com, investopedia.com and hrexecutive.com. Twelve run the other way, cited by Gemini and by neither Search surface: techradar.com, allaboutcookies.org, useboomerang.com, gitnux.org, safetydetectives.com, safewise.com, unbiased.com, time.com, zipdo.co, goodhousekeeping.com, security.org and medicalbillersandcoders.com. Even the two Search surfaces diverge: 29 top-250 domains are cited by AI Mode and never by AI Overviews, and seven the reverse.
Google's documentation describes the control that separates the first list from the second. Google-Extended is a robots.txt token that governs whether crawled content may be used for training Gemini models and "for grounding" in Gemini Apps; it "does not impact a site's inclusion in Google Search." Of the fourteen Gemini-absent domains, twelve had a readable robots.txt on September 18, and eight of those twelve disallow Google-Extended outright: Yahoo, NerdWallet (which blocks nothing else), TechCrunch, The New York Times, CNBC, Instagram, Investopedia and U.S. News. Of the twelve domains on the reverse list, eleven had a readable file and none of them disallows Google-Extended; one, security.org, allows it explicitly. For the Search surfaces, Google's AI features guidance sets one condition, that a page be indexed and eligible for a snippet in Search, and notes that "AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary."
TechRadar is the case that no file explains. It carries no AI-specific robots rules, is cited by ChatGPT, Claude, Gemini and Perplexity at an A-grade 2.84%, and appears in neither Google Search surface across 15,782 runs. Whatever governs that is on Google's side of the line, and the Index can only record it.
Perplexity: the widest engine, and where it gets Reddit
Perplexity cites Reddit and Facebook, whose files disallow every agent they do not name, and The New York Times and PCMag, which disallow PerplexityBot and Perplexity-User by name. Perplexity's crawler documentation distinguishes PerplexityBot, which builds search results and honours robots.txt, from Perplexity-User, which fetches a page when a user asks a question and "generally ignores robots.txt rules." Reddit's October 2025 complaint alleges a third route: that Perplexity and its co-defendants obtained Reddit content by scraping Google's search results, which Reddit tested by posting content visible only in Google's results and watching it surface in Perplexity's answers within hours. Perplexity denied wrongdoing. The Index does not measure which route produced a citation; it records that Perplexity is the one engine of six with Reddit present and the widest source layer at every tier.
What other panels see, and why they differ
Two other measurements are worth setting beside this one, because they disagree with it in an informative way. Semrush and Kevin Indig ran 100 buyer-journey prompts through GPT-5.2 in minimal and high reasoning and found Reddit cited in both modes, with Reddit and other user-generated sources losing "roughly half their share of citations" when reasoning is on. Loudmink's vendor panel of 20 B2B SaaS queries across five engines put Reddit at 3 to 5% of ChatGPT's citation URLs, zero for Claude, and about 2% for Perplexity through May 2026 before what it describes as a reversal to about 90% of answers in June.
The Index's ChatGPT runs show zero Reddit citations over 125 days. The difference is the panel. The Index asks 912 buyer questions across 24 subject categories in six question shapes, from the same fixed basket every day, and reports a domain present only if it was cited at least once in a run. A panel of twenty queries with 25 named brands, or 100 prompts in consumer categories, samples a different slice of what an engine will cite. Both can be right about their own panels. The claim that survives all three is the one on Claude: no measurement reviewed here has recorded Claude citing Reddit.
What this changes for publishers and brands
The engine is the unit of measurement. A source that ranks first across the market can be structurally absent from a third of the engines, and a robots.txt written for training opt-outs is now a per-engine distribution decision: a Google-Extended block removes a site from Gemini's grounding while leaving it eligible for AI Overviews, and a wildcard block that does not name OAI-SearchBot or Claude-SearchBot removes it from two engines' search answers even if GPTBot and ClaudeBot are allowed. A licensing agreement, on the evidence of Reddit and The Wall Street Journal in this panel, guarantees access, not citation. Anyone measuring their own presence should read it engine by engine, on a fixed panel, over a window long enough to see whether a split holds. This one has held for two months.
Change log
- Measured change since July 25, 2026. Reddit's overall citation rate rose from 11.55% (1,069 of 9,258 runs) to 12.75% (2,012 of 15,782 runs). Engine presence unchanged at four of six; ChatGPT and Claude absent in both reads.
- New in this read. Per-engine width across the universe (Perplexity 9,497 distinct domains to AI Overviews 2,279); the Gemini-versus-Search split and its match to Google-Extended; the nine head domains absent from both ChatGPT and Claude.
- Watch next. Whether the OpenAI Reddit partnership becomes visible in ChatGPT citations on this panel; the Reddit v. Perplexity docket; TechRadar's absence from both Google Search surfaces at the next release.
Method note
Source: the public Machine Relations Index JSON, release mri_score_v2.0+2026-09-18+8fa38e54dd0a, window 2026-05-10 to 2026-09-18, 125 observed days, 15,782 answer runs, 124,397 citation events, 22,179 cited domains, six engines, evidence floor of 10 runs across 7 dates for any published segment rate. Rank here is the overall citation-rate order across all measured runs. Per-engine figures are presence (cited at least once in the window) and never a rate; the public view carries no per-engine rate. Domain profiles for reddit.com, techradar.com, facebook.com and nytimes.com are published by the Index; the week's full read is the Index Report for September 18. Robots.txt files were read on September 18, 2026, and change without notice. Correlation between a robots rule, an agreement or a lawsuit and an engine's citation behaviour is reported as correlation. Paralax's earlier note on Reddit's search-referral warning covers the traffic side of the same relationship.
FAQ
Which AI engines cite Reddit? In the Index's September 18 release, Gemini, Google AI Mode, Google AI Overviews and Perplexity. ChatGPT and Claude cited reddit.com in zero of 15,782 runs.
Why would ChatGPT not cite Reddit when OpenAI has a Reddit agreement? The Index cannot say why. What is documented is that Reddit's robots.txt disallows every agent including OAI-SearchBot, and OpenAI states that sites opted out of OAI-SearchBot are not shown in ChatGPT search answers. Other panels, run on different prompts, do record ChatGPT citing Reddit.
Does blocking Google-Extended remove a site from AI Overviews? Not according to Google's documentation, which says the token controls Gemini training and grounding and does not affect Search inclusion. The Index's data is consistent with that: eight top-250 domains that block Google-Extended are cited by both Search surfaces and never by Gemini.
Is a per-engine citation rate available? No. The public Index records which engines cited a domain at least once in the window. Rates are published per domain and per segment across all engines.
Which engine cites the most sources? Perplexity, on every measure in this release: 9,497 distinct domains, 4,246 cited by no other engine, and presence on 967 of the top 1,000.