The 42-point “ranking premium” in AI citations is a mirage — and a new causal audit shows why your dashboards are lying
Hi everyone, this is Neo.
If you work on GEO, you’ve seen the chart. Some AI visibility platform plots a tidy downward curve — sources at the top of a search call get cited most often, sources further down get cited less — followed by the confident conclusion: go win the top spot.
On September 14, a preprint landed on arXiv titled CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search, by Sriram Selvam and Anneswa Ghosh. What it did is simple, and what it found is bad news for every dashboard in the space.
That curve is mostly not about position. It’s mostly about search engines putting the more relevant page first.
Here’s the design, the three findings, and what to change in your own measurement.
The setup: an experiment that inverts the obvious reading
Start with the number everyone quotes. In their data, documents first exposed at rank 1 in a search call were cited 85.1% of the time. At rank 5, 42.8%. A 42.3-point gap.
Anyone would read that as “position matters enormously.” The paper points out the flaw: position and quality are bound together. Search engines rank more relevant, higher-quality pages earlier, and in an agentic setting the agent chose the queries that produced those rankings. So the chart is actually measuring “good content vs. average content.”
So they did something smarter:
- Freeze real conversations. A tool-using GPT-5.4 agent answered 130 everyday questions (10 topics × 5 phrasings, from full questions to terse search-style input). It had to run its own web searches — five Exa results per call. Every message, result object and retry was archived: 129 conversations, 346 completed search calls, 1,490 unique documents.
- Look for competing pairs, not rankings. A separate evaluator judged whether a document supports a specific answer fact — blind to rank, title, domain, answer and citations, and required a verbatim quote with exact string matching. Paraphrased or missing quotes were downgraded. Of 1,465 reviewed documents, 1,115 (76.1%) were answer-bearing. Then they selected competing pairs: two documents from the same search call, each independently supporting the same fact — meaning either one could legitimately be cited. At least 1,894 candidate pairs existed; 113 were frozen, and a blinded human check confirmed 103 of those (91.2%) as genuine same-proposition competitions.
- Hash-verified replay, one variable at a time. Change only the order of the two documents in the array, or swap the target’s rendering (plain prose vs. headings, short paragraphs and lists/a table) — everything else byte-identical, only the final answer regenerated. Four cells per pair (order × rendering): 452 trials, zero integrity failures.
The elegance is that this turns “who got cited” from a correctness question into an allocation question. Both documents support the same fact. The engine’s choice is pure credit-splitting.
Finding 1: the observed position gap is five times the controlled effect
| Measurement | Result | Statistical read |
|---|---|---|
| Observational gap (rank 1 vs rank 5) | 85.1% → 42.8%, 42.3 points | Descriptive; confounded with relevance and quality |
| Controlled order effect (main replay) | +7.9 points (95% CI [+1.1, +14.9]) | Holm p = 0.350 — not significant |
| Held-out confirmation (56 pairs, order swapped only) | 0.0 points (95% CI [−5.4, +5.4]) | Point estimate is exactly zero |
The paper is careful: message position can move attribution in some frozen transcripts, especially with large displacements. But average rank effects inferred from observational gradients are not reliable.
In plain language: the “position premium” on your dashboard is not a decision input. It credits the retrieval engine’s ordering and your content quality to a single variable called position.
Finding 2: the one effect that survived correction is credit concentration
This is the genuinely useful part for content teams.
| Effect | Size | 95% CI | Significance |
|---|---|---|---|
| Target’s citation count (structured vs. prose) | +0.50 markers per answer | [+0.20, +0.84] | Holm p = 0.033 (only corrected secondary result to survive) |
| Target cited at least once (pre-registered primary) | +4.5 points | [−1.4, +10.4] | p = 0.168 — not significant |
| Matched competitor’s citations | −0.02 markers | [−0.28, +0.24] | Essentially unchanged |
| Answer-wide citation budget | −0.49 markers | crosses zero | No inflation — the pie didn’t grow |
| Share of added credit attaching to shared evidence | 48.6% (+0.25 markers, p = 0.023) | — | Post-hoc audit |
Three details, or this gets misread as “add bullet points, win citations”:
First, what changed is allocation, not admission. Structured formatting did not significantly raise the chance the target gets cited at all (+4.5 points, CI crossing zero). What it raised is how much credit that page collects once it’s in the citation set — on evidence both pages support. Total answer citations didn’t grow, and the competitor didn’t lose. Nobody made the cake bigger; one name got written on more slices.
Second, about half the added credit lands on evidence both documents support. Structure doesn’t help you get your exclusive content into the answer. It helps your name become the one attached to a fact someone else also has.
Third, a strict ablation took most of the shine off. The researchers also ran a word-preserving reformat — the exact same sentence sequence relaid as one sentence per list row, no model involved. Across all 113 pairs, incidence rose +6.7 points (p = 0.028). On the 30 families used for repeatability, it came out −1.7 first, then −8.3 on a fresh generation: −5.0 points averaged (p = 0.22).
So there is no stable list-marker mechanism. The paper’s own framing: visible presentation can redistribute credit, but the pathway is transcript-dependent — not a recipe you can copy.
Finding 3: citation measurement has a measurable noise floor
This is the page I want every GEO report owner to read.
They regenerated 120 cells with identical inputs:
- The binary “was the target cited?” decision agreed only 85.0% of the time — roughly one flip in seven;
- Exact cited-pair sets agreed 78.3%;
- Exact target citation counts agreed 61.7%;
- A method-of-moments decomposition attributes about 45% of single-draw family-effect variance to decoding randomness;
- Only 17 of 30 families showed consistent family effects.
The paper puts it bluntly: citation studies reporting a single generation inherit a quantifiable noise floor, not merely unspecified stochasticity. Repeat k times and you shrink that noise component by roughly 1/k.
The difference from “we asked the question twice” is that here the question, the phrasing, and the inputs are identical.
The boundaries the paper draws (keep these handy)
If you’re going to cite this at work, bring the caveats too:
- One retrieval provider (Exa, five results per call) and one primary model (GPT-5.4). Nothing here generalizes to all AI search.
- It only altered the text the model reads — downstream of crawling, retrieval and ranking. So it does not show that editing your page raises AI citations. It shows that once text is in the model’s context, presentation shifts credit.
- Both renderings were AI rewrites (with a seven-dimension fidelity gate), so the contrast is a rewrite package, not pure formatting.
- The primary incidence hypothesis was underpowered — the design reliably detects effects around 8.5 points or larger, so +4.5 is unresolved, not proof of zero.
- Cross-model replay (Grok 4.3) leaned the same way but only 44.7% of its first responses satisfied the citation contract; after repair the effect was +5.6 points. Directional support, not magnitude replication.
And the single most important line in the discussion section, in the authors’ own words:
“This is an attribution-sensitivity warning, not an optimization tactic.”
Two cross-checks worth knowing
The paper also touches two popular industry claims, in the spirit of its own method:
- How unstable is a single run? SparkToro reported in January that giving ChatGPT and Google’s AI Overviews the same prompt repeatedly produced the same brand list less than 1% of the time.
- Does JSON-LD help? An Ahrefs report in May found pages cited by AI were about three times as likely to carry JSON-LD schema. The report could not show — and doesn’t claim — that adding schema increases citations. That’s correlation. The whole point of this preprint is that a correlation only becomes causal after you change that one variable in isolation.
Neo’s take: none of this means structure or schema is useless. It means almost every GEO number in circulation has not yet passed the controlled-variable test. The industry is at a stage where its measurement methods trail its commercial promises by a wide margin.
Six changes to make in your own reporting
1. Run every cell at least three times. With a one-in-seven flip rate as your baseline, a single run has no authority. A brand “losing a citation” on one prompt shouldn’t reach a slide deck before repeat sampling.
2. Change how you phrase results. Not “our AI citation rate rose from 32% to 41%” — instead, “across three samples, our citation rate ranged X–Y.” One is a noisy sample. The other is a claim you can defend.
3. Stop paying for position. Position is the retrieval engine’s internal ordering. You don’t control it, and it’s heavily collinear with content quality, so “earning the top slot” is usually an echo of relevance you already had. What you do control is extractability (can the content be read cleanly?) and evidence density (how many independently citable facts live in a paragraph?).
4. Aim structure at the right target. Not “get listed in the citations” — but “be the source credited for the fact everyone shares.” That’s the operational takeaway of the paper. Write the industry-consensus paragraph with a clear heading, tight points and a table, and you may not get mentioned more often, but the credit is more likely to land on your version.
5. Run small, comparable A/B tests. Change one thing at a time — e.g. convert a single section to structured format — run it repeatedly, and look for a stable direction. Not ten changes at once followed by a look at the aggregate line. That’s exactly what the researchers did: hold everything else fixed first.
6. Build a monthly repeat-sampling table. Fix 20–30 prompts, three engines, three samples per prompt. The value isn’t any single score. It’s that you can recognize a real change the moment it exceeds your noise band.
Final word
This paper hands out no quick wins. It hands out a calibration ruler: a 42.3-point ranking premium that goes to zero under control, one surviving effect (structure concentrating credit, +0.5 citation markers), and a 15% flip rate baked into any single generated answer.
The good news for site owners: you don’t control AI ordering, but you do control whether your content is clean, factual, easy to extract. And per this paper, at least the credit-concentration effect sits on your side of the line.
The bad news: if you keep making calls off a red-and-green grid built on single runs, you’ll spend next quarter optimizing a variable that isn’t there.