AI Engines Do Not Swap Pages, They Swap Whole Sources: 89 Days of Per-Engine Citation Churn
A first-party 89-day panel of 35 queries across five AI retrieval surfaces measures day-over-day citation churn per engine, then tests the category's standard explanation for it. The two-tier model — sticky domains, rotating pages — does not appear. When an engine keeps a domain, it cites the exact same URL 72% to 91% of the time.
If an AI engine cited your page yesterday, the odds it still cites you today depend almost entirely on which engine it was — and when you lose the citation, you usually lose the whole domain, not just that page.
We ran the citation-volatility study on our own panel: 35 fixed queries, five AI retrieval surfaces, once a day, 89 logged runs between June 24 and September 22, 2026. That is 11,539 engine-query observations carrying 58,855 citation links. Every answer and its cited source list is stored, so the day-over-day comparison is a set comparison against our own evidence, not a recollection.
Two things came out of it. The first is a per-engine churn table that puts a number on how unstable each surface is. The second is more useful: the standard industry explanation for that churn does not survive contact with the data.
What we measured, and what it is not
This matters more than the numbers, so it goes first.
Four of the five surfaces are each vendor's API with its search tool enabled, not the consumer chat product:
| Surface | What we actually query | Model held constant over the window |
|---|---|---|
| Perplexity | sonar |
Yes, all 89 days |
| ChatGPT | OpenAI web_search tool |
No — two model changes |
| Claude | Anthropic web_search tool on Haiku 4.5 |
Yes, all 89 days |
| Gemini | Grounding with Google Search | No — five model versions |
| Google AI Mode | The consumer AI Mode surface, captured by scraper | Not applicable |
Only Google AI Mode is the consumer surface. Calling the other four "ChatGPT" or "Gemini" is shorthand for a retrieval layer, and anyone comparing these numbers to a study run in logged-out consumer sessions is comparing two different populations. Our average cited set is 3.4 to 5.2 domains per answer; a 530,875-citation consumer-session study by GetMentions reports 8.0 to 40.2 distinct domains per query. Smaller cited sets mechanically produce different churn arithmetic. We are not claiming our magnitudes transfer to theirs.
The query panel is also narrow and deliberate: 35 registered queries about AI visibility, AI citation and Machine Relations — one category, the one we work in. Twenty-six are core-bucket terms. This is a category panel, not a sample of the web.
And it is one sample per engine per query per day. Some share of what we call "churn" is single-sample answer variance rather than genuine change, and we cannot separate the two without repeated same-day sampling, which this panel does not do. What the number does describe exactly is what any once-a-day tracker reports to a brand — because it is the same measurement.
Day-over-day churn, per engine
Churn is the share of yesterday's cited domains that are gone from today's answer to the same query.
| Engine | Day-pair observations | Avg domains cited | Domain churn, D+1 | URL churn, D+1 |
|---|---|---|---|---|
| Perplexity | 2,694 | 5.1 | 22.9% | 24.3% |
| Claude | 1,997 | 3.8 | 38.0% | 40.5% |
| ChatGPT | 2,220 | 3.4 | 38.7% | 44.1% |
| Google AI Mode | 1,251 | 5.0 | 55.0% | 59.1% |
| Gemini | 1,969 | 5.2 | 57.6% | 62.5% |
Perplexity replaces roughly one source in four overnight. Gemini and Google AI Mode replace more than half. The spread between the steadiest and the least steady surface is 34.7 percentage points — larger than the difference between being cited and not being cited on most individual queries.
Stretch the window and the ordering holds but the gap closes:
| Engine | Domains retained after 1 day | After 7 days | After 28 days |
|---|---|---|---|
| Perplexity | 77.1% | 62.7% | 44.2% |
| Claude | 62.0% | 57.9% | 50.5% |
| ChatGPT | 61.3% | 54.7% | 45.9% |
| Google AI Mode | 45.0% | 37.4% | 28.8% |
| Gemini | 42.4% | 37.6% | 31.8% |
Claude is the interesting column. It starts less stable than Perplexity by 15 points and ends more stable by 6. Perplexity's day-one advantage is partly short-term stickiness that erodes; what survives a day on Claude mostly survives a month. If you are buying against a 30-day reporting cycle, those two engines rank in the opposite order from how they rank on a daily dashboard.
The confound we had to rule out first
Gemini changed model version five times inside the window — 2.5-flash, 3.5, 3.6, 3.7, 3.8 — and ChatGPT twice, while Perplexity and Claude held one model for all 89 days. The churn ranking lines up suspiciously well with that: the two surfaces we never changed are the two steadiest.
So the honest read is that "Gemini is volatile" might mean "we swapped the model underneath Gemini five times."
We tested it by splitting Gemini's day pairs into those where both days ran the same model version and those that straddle a version change:
| Gemini day pairs | Count | Domain churn |
|---|---|---|
| Same model version on both days | 1,875 | 57.4% |
| Across a model version change | 94 | 61.0% |
A version change costs about 3.6 additional points of churn, and only 94 of 1,969 pairs cross one. Gemini's instability is a property of the retrieval surface, not of our rotation. The finding survives its strongest alternative explanation, which is the only reason it is published.
No vendor volatility study we have read discloses the model version behind its captures, which means none of them can run this check.
The two-tier model does not appear here
The category has a standard reconciliation for citation volatility, and we publish it ourselves in a research synthesis of five independent studies and a companion evidence page. It goes: domains are the structural anchor and stay in the cited set; the specific page beneath the domain is the replaceable unit and rotates constantly. Domain-level stability measured at 96.8% week over week in one study; page-level one-time appearance at 44% in another; both correct, different layers. Sistrix reports the same shape from a third angle, Trakkr's longitudinal work and Profound's drift tracking sit alongside it, and the Writesonic 23-million-source study is the usual source of the page-level half.
That model makes a sharp, testable prediction: URL-level churn should be dramatically higher than domain-level churn, because pages rotate under stable domains.
On our panel the gap is 1.4 to 5.4 percentage points.
| Engine | Domain churn | URL churn | Gap |
|---|---|---|---|
| Perplexity | 22.9% | 24.3% | 1.4 pts |
| Claude | 38.0% | 40.5% | 2.4 pts |
| Google AI Mode | 55.0% | 59.1% | 4.1 pts |
| Gemini | 57.6% | 62.5% | 4.9 pts |
| ChatGPT | 38.7% | 44.1% | 5.4 pts |
We measured the mechanism directly as well. Take every case where a domain was cited for a query on one day and still cited for the same query the next day, and ask what happened to the specific URLs:
| Engine | Domain-days retained | Exact same URL set | Partial overlap | Full page swap |
|---|---|---|---|---|
| Perplexity | 10,665 | 90.9% | 7.0% | 2.1% |
| Claude | 4,750 | 83.9% | 12.7% | 3.4% |
| ChatGPT | 4,595 | 82.7% | 11.2% | 6.1% |
| Google AI Mode | 2,811 | 80.0% | 13.5% | 6.5% |
| Gemini | 4,375 | 72.0% | 17.6% | 10.4% |
When an engine keeps a domain, it cites the identical URL set 72% to 91% of the time. It swaps to entirely different pages on that domain between 2.1% and 10.4% of the time.
The domain and the page are not two layers here. They are one unit that holds or falls together. Being "on a domain AI trusts" bought nothing on this panel that the specific page did not already have.
We are not calling the two-tier studies wrong. They run far larger panels on consumer surfaces, and the structure they describe may be exactly what happens at that scale. But the reconciliation is repeated across the category as settled physics, and on the one panel where we can inspect every cited URL against its own domain on consecutive days, it is absent. Nobody publishing the two-tier claim publishes this test.
Losing a citation is usually not losing a citation
The other operational finding. Take every domain that was cited one day and gone the next, and follow it:
| Engine | Drop-outs | Back within 7 days | Back within 28 days | Still gone at 28 days |
|---|---|---|---|---|
| ChatGPT | 3,151 | 61.3% | 74.9% | 25.1% |
| Claude | 3,194 | 57.6% | 72.8% | 27.2% |
| Perplexity | 3,185 | 56.5% | 68.3% | 31.7% |
| Gemini | 6,399 | 44.4% | 58.5% | 41.5% |
| Google AI Mode | 3,962 | 36.5% | 47.1% | 52.9% |
Between 47.1% and 74.9% of dropped domains are back inside a month. Most disappearances are oscillation, not displacement. A daily alert that says you lost a citation is, on four of these five surfaces, more likely to be describing a source that returns than a source that is gone — and acting on it by rewriting a page is a response to noise.
Google AI Mode is the exception worth flagging: a majority of its drop-outs do not come back within 28 days. It is also the surface whose capture coverage degraded during our window, so read that row as the weakest in the table.
The core is small and it is where the durability is
Across queries observed on at least 30 days, count how often each domain appears:
| Engine | Distinct domains cited | Cited on ≥50% of days | Share of all citation slots they hold | Cited exactly once in 89 days |
|---|---|---|---|---|
| Perplexity | 948 | 101 (10.7%) | 44.5% | 23.9% |
| ChatGPT | 668 | 58 (8.7%) | 43.3% | 28.6% |
| Claude | 789 | 67 (8.5%) | 49.6% | 40.3% |
| Google AI Mode | 1,822 | 60 (3.3%) | 30.3% | 64.5% |
| Gemini | 2,297 | 56 (2.4%) | 27.5% | 51.4% |
Every surface has the same shape: a small core carrying a third to a half of all citations, and a long tail of domains that appear once and never return. On Google AI Mode, 64.5% of every domain it ever cited in 89 days appeared on exactly one day. On Gemini, 51.4%.
This is where the engine ordering becomes actionable. A single citation on Gemini or Google AI Mode is, at base rate, most likely a one-day event. A single citation on Perplexity or ChatGPT is much more likely to be the start of a run. The same screenshot means two different things depending on which engine produced it.
It is also the one place our panel and the consumer-session panel agree. GetMentions, running 2,398 consumer prompts over seven days with a completely different capture method, reports one-day-only source shares of 33.4% for Perplexity, 53.6% for ChatGPT, 66.2% for Gemini and 67.5% for Google AI Mode. Different panel, different surface, different window, different absolute numbers — same rank order as ours. When two independent measurements disagree on magnitude and agree on ordering, the ordering is the part you can act on.
What to do with this
Stop reading daily citation movement as performance. On four of five surfaces the majority of losses reverse within a month. Set your reporting window to 28 days or do not report movement at all.
Pick the engine your measurement contract is denominated in. Perplexity is the steadiest surface on a one-day view and Claude on a 30-day one. A vendor showing you a stability number without naming the engine and the window is showing you an average of two different behaviours.
Stop buying domain authority as a citation strategy on the strength of the two-tier story. On this panel, having a page on a frequently-cited domain did not get that page cited when a sibling page was dropped. The engines swapped the whole source.
Ask any AI-visibility vendor what model version they captured. Five model versions passed under one of these engines in 89 days. A volatility number that does not pin the model cannot tell you whether it measured the web or measured the vendor's own release cycle.
The measurement that produced this page is the same one behind our work on per-engine citation support rates and the AuthorityTech measurement practice that uses it, and it sits beside the Machine Relations Index as the query-side companion to the index's source-side view. The panel runs daily and the next read on it will say whether these ratios hold as the surfaces change underneath them.
Method
Panel: 35 registered queries, 89 logged runs between June 24 and September 22, 2026, five surfaces, 11,539 engine-query observations, 58,855 citation links. The 89 runs cover 85 distinct dates inside that 91-day span: six dates have no run (August 6, August 15, August 31, September 3, September 6, September 9) and four dates carry two runs. Consecutive-run pairs are therefore not always consecutive calendar days.
Correction, September 22, 2026. This page originally described the panel as "89 consecutive daily runs". It is 89 logged runs over 85 distinct dates in a 91-day span; six dates are missing and four carry two runs, so a small share of the day-over-day pairs span a two-day gap or sit inside a single date. Recomputing pooled one-day domain retention over strictly adjacent calendar dates only moves each engine by at most 0.6 percentage points, so every conclusion above stands; the description of the panel was wrong and has been corrected here rather than quietly amended. One capture per engine per query per day. Citation lists are taken from each provider's returned source set, stored per run with the answer text.
Domain normalisation strips www. and lowercases the host. URL normalisation drops query strings, fragments and trailing slashes. Churn at lag n compares the cited set for a query on day D with the cited set for the same query on day D+n, counting only pairs where both days produced a non-empty cited set for that engine; pairs where an engine returned nothing are excluded rather than scored as total churn. Re-entry follows each dropped domain forward through that query's subsequent observation days within a 28-day horizon.
Limits, stated once more because they bound every number above: this is a single-category query panel, one sample per day, and four of the five surfaces are vendor APIs with search tools rather than the consumer products those names usually mean. Google AI Mode's capture coverage declined during the window and its rows are the least reliable in every table. None of these figures should be read as the behaviour of AI search in general; they are the behaviour of five specific retrieval surfaces on one panel that we can inspect completely.