Reddit Is the #1 Cited Domain in the Machine Relations Index — and Two of Six Engines Have Never Cited It
Reddit ranks #1 of 22,320 domains in the Machine Relations Index, but its engine breadth is 4 of 6: ChatGPT and Claude show zero citations in the same window.
Reddit is the single highest-ranked domain in the Machine Relations Index: #1 of 22,320 observed domains, cited in 2,016 of 16,039 monitored answer runs, a 12.57% citation rate, Confidence A. Read that sentence on its own and it says Reddit is the source AI engines trust most.
Read the domain's own profile page and a second number sits right below the first: engine breadth 4 of 6. The six engines the Index monitors are ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity. Reddit's citations in this window come from four of them — gemini, google_ai_mode, google_ai_overviews and perplexity. ChatGPT and Claude do not appear in Reddit's engine list at all, not at a low rate, not occasionally — zero observed citations from either engine across the full 2026-05-10 to 2026-09-20 window.
That is the finding this piece exists to state plainly, because the rank alone hides it. A leaderboard position is a market-position claim; an engine-breadth field is a crawl-coverage claim, and they answer different questions. Reddit's #1 rank is real and it is Confidence A — a high-volume, high-confidence number. But "the source AI engines cite most" is not the same claim as "the source every AI engine cites," and reporting the first as if it were the second is exactly the leaderboard error worth naming: a rank without its per-engine breakdown reads as consensus when it may be one or two engines' crawl coverage carrying the total.
Read live, release mri_score_v2.0+2026-09-20+47973f373a20
Every figure below is read from machinerelations.ai/index/domains/reddit.com and the release's top-level dataset page, both fetched live by plain HTTP on 2026-09-20, no sandbox required. Release mri_score_v2.0+2026-09-20+47973f373a20, window 2026-05-10 to 2026-09-20 (127 days), 16,039 monitored answer runs, 22,320 cited source domains, six answer engines.
| Domain | Overall standing | Citation rate | Cited runs | Engine breadth | Confidence |
|---|---|---|---|---|---|
| reddit.com | #1 of 22,320 | 12.57% | 2,016 / 16,039 | 4 of 6 (missing chatgpt, claude) | A |
| youtube.com | #2 of 22,320 | 8.95% | 1,435 / 16,039 | 6 of 6 | A |
| linkedin.com | #3 of 22,320 | 6.79% | 1,089 / 16,039 | 6 of 6 | A |
| nih.gov | #6 of 22,320 | 3.04% | 487 / 16,039 | 6 of 6 | A |
| g2.com | #8 of 22,320 | 2.84% | 456 / 16,039 | 6 of 6 | A |
| gartner.com | #9 of 22,320 | 2.84% | 456 / 16,039 | 6 of 6 | A |
| microsoft.com | #12 of 22,320 | 1.90% | 305 / 16,039 | 6 of 6 | A |
| healthline.com | #15 of 22,320 | 1.86% | 298 / 16,039 | 6 of 6 | B |
| medium.com | #1 of 1,230 editorial | 5.04% | 808 / 16,039 | 6 of 6 | A |
| forbes.com | #2 of 1,230 editorial | 4.15% | 665 / 16,039 | 6 of 6 | A |
| businesswire.com | #74 of 22,320 | 0.61% | 98 / 16,039 | 6 of 6 | C |
Every other domain that clears Confidence A or B standing in the top 15 of the full index carries full six-engine breadth. Reddit, at #1, is the outlier: the highest-ranked domain in the entire measured universe, and the only one of this group with two engines showing no observed citations at all.
Thinner, low-confidence domains show narrower breadth too, but for the expected reason — small sample. Instagram sits at #58 on 114 cited runs with 3 of 6 engines. Quora sits at #164 on 60 cited runs with 4 of 6, Confidence C. Stack Overflow and TikTok are still in the "collecting" tier with 3 and 4 of 6 respectively, on fewer than 30 cited runs each. Reddit does not fit that pattern: 2,016 cited runs, Confidence A, the top rank in the index, and still two engines at zero.
Why the gap does not resolve itself by counting harder
The Index's own rule is the reason this needs the domain-profile field and not the leaderboard alone: category leaderboard pages carry rank and rate, not a per-engine column. The only place engine breadth is published is the domain's own profile. A researcher who reads Reddit's #1 leaderboard position without opening its profile page has no way to know the total is built from four engines, not six — the exact shape of error worth naming: a rank presented without the breadth that supports it reads as agreement it does not have.
Reddit's public record supplies plausible, non-exclusive context for the split, offered here as correlation rather than a causal claim the citation data itself does not measure. Reddit signed content-licensing partnerships with both Google and OpenAI in 2024 — the Google deal reported at roughly $60 million a year, confirmed by TechCrunch the same day, and the OpenAI partnership granting ChatGPT access to Reddit's Data API. A signed OpenAI deal exists, and ChatGPT still shows zero citations of Reddit in this window — the licensing relationship and the observed citation behavior do not move together, which is itself worth naming rather than assuming a licensing deal predicts citation volume. Reddit's relationship with Anthropic is adversarial rather than contractual: Reddit sued Anthropic in June 2025 alleging Claude was trained on scraped Reddit content without a license, and in October 2025 sued Perplexity and three data-scraping intermediaries over unlicensed access, a filing also reported by Gigazine — yet Perplexity is one of the four engines the Index observes citing Reddit in this window, and Claude, the other litigation target, is not. A Cloudflare-documented forensic report on Perplexity's crawling behavior sits in the same public record, and Reuters covered the market reaction to the OpenAI deal separately. None of this explains the ChatGPT/Claude split cleanly; it says the legal and commercial relationship between Reddit and each engine operator is not a reliable predictor of which engines cite it, which is a reason to read the engine-breadth field directly rather than infer it from deal announcements.
What the crawler documentation actually controls
Citation behavior is not the same thing robots.txt governs, and the two should not be conflated. Google's own crawler documentation states that the Google-Extended token controls whether crawled content may be used to train future Gemini models and for grounding in Gemini Apps — not inclusion in Google Search, and not a ranking signal. OpenAI documents GPTBot as the token that controls training-data access, separate from the browsing agents that can still fetch a page to ground a specific answer. Anthropic documents its three crawlers — ClaudeBot, Claude-User and Claude-SearchBot — and what disabling each one costs a publisher, and Perplexity documents which bots it operates under similar logic: an access-control token is a publisher's stated policy, not proof of why any specific citation appeared or did not. Reading a domain's robots.txt is evidence of what the publisher has said engines may do, never the mechanism by which a citation is observed or withheld.
The read for measurement, not for optimization
This is not a ranking-manipulation story and it does not carry a recommendation to chase Reddit citations or avoid them. It is a measurement note: the single highest-ranked source in a 22,320-domain index can still be a four-engine phenomenon, and the only way to know that is to open the domain's own profile rather than stop at the leaderboard. Every top-15 domain in this data except Reddit carries full six-engine breadth; Reddit does not, and its rank does not say so on its own. Any citation-rate leaderboard — ours or anyone else's — is a crawl-coverage claim unless engine breadth rides beside the rank.
Method
Every domain figure in this piece was read live on 2026-09-20 by plain HTTP GET against machinerelations.ai/index/domains/<domain>.md for reddit.com, youtube.com, linkedin.com, nih.gov, g2.com, gartner.com, microsoft.com, healthline.com, medium.com, forbes.com, businesswire.com, quora.com, x.com, wikipedia.org, stackoverflow.com, tiktok.com, instagram.com and pinterest.com, plus the release summary at machinerelations.ai/index, all against release mri_score_v2.0+2026-09-20+47973f373a20. No sandbox was used; these are the public per-domain routes, not the 59 MB aggregate JSON. "Engine breadth" here means the count of the six monitored engines with at least one observed citation of the domain in the window — a presence count, never a per-engine rate, since the public dataset does not publish rate per engine per domain. This piece cites none of the four adjudicated study families and introduces no rolling-window figure of our own; every third-party fact carries its own dated source above.