A September 7, 2026 arXiv preprint by Elisha Bajemon and Andre-Louis Rochet tested a question GEO scorecards often skip: does a page-only content score predict which source a modern generative engine will cite? In the original seven-engine outcome set, the pooled within-query Spearman correlation was only 0.114 across 777 evaluable query-engine groups.
The paper also re-tested three original GEO interventions: quotations, statistics, and source citations. None produced a positive pooled effect. Two confidence intervals included zero; cite-sources was negative with an interval excluding zero.
This does not mean “GEO does not work.” This query-blind score weakly predicted citation ordering, while the three historical interventions did not transfer positively. Query-conditioned models reached about 0.37 to 0.38.
Direct answer — Do GEO scores predict AI citations?
In this preprint, a deterministic content-only GEO score had a weak positive relationship with citation visibility: within-query Spearman rho was 0.114 on seven original engine arms and 0.118 in a three-model GPT-5.x replication. That is not zero, and it is not 11.4% accuracy. It means the score only weakly preserved the engines’ ordering of five candidate sources for the same query.
Key Takeaways
- The work is an arXiv preprint, not peer reviewed.
- The fixed score reached rho = 0.114 across 777 evaluable groups; three GPT-5.x arms replicated at 0.118 across 450 groups.
- Quotation, statistics, and cite-sources edits produced 0 of 3 positive pooled effects.
- The ranking data contain 229 unique queries and 1,269 query-engine groups, with five candidate sources per group.
- Correcting query leakage cut a fitted-model result from 0.5724 to 0.3625.
What the Study Actually Tested
The study separates two objects. First is a deterministic 0–100, query-blind score built from 11 text features: entropy, information and entity density, semantic coherence, self-containment, citation F1, statistic density, MMR, nDCG, redundancy, and quotable density. Second is a query-conditioned predictor that adds query-source relevance and asks how well it can reproduce the engine’s ordering of five candidate sources.
The denominators differ. The 500-source benchmark tests score-gaming resistance, not rho 0.114. Citation-ranking data contain 229 unique queries and 1,269 query-engine groups. The original seven-arm analysis has 777 evaluable groups; GPT-5.5, GPT-5.4, and GPT-5.4-mini add 150 each.
The outcome is a PAWC-like visibility share in generated answers, not Google rank, traffic, leads, or revenue. Keep those layers separate in AI search visibility measurement.
Three Popular GEO Tactics Showed No Positive Pooled Lift
| GEO claim / tactic | What was actually tested | Observed result | Evidence supports | Still unknown |
|---|---|---|---|---|
| A page-only GEO score predicts citations | Query-blind score versus within-query source visibility | rho 0.114, n=777; GPT-5.x rho 0.118, n=450 | Weak positive ordering relationship | Whether commercial GEO scores predict live-search citations |
| Add quotations | Paired, volume-controlled quotation edit | -0.325 pp; 95% CI -0.841 to 0.171; n=1,531 | No positive pooled effect detected | Other styles, domains, prompts, or retrieval systems |
| Add statistics | Paired, volume-controlled statistics edit | -0.276 pp; 95% CI -0.970 to 0.425; n=1,087 | No positive pooled effect detected | Whether original or highly query-relevant data differ |
| Add source citations | Paired, volume-controlled cite-sources edit | -0.793 pp; 95% CI -1.533 to -0.138; n=1,087 | No positive effect; pooled estimate was negative | Whether source quality or implementation changes the result |
| Query context matters | Query-aware models on query-disjoint folds | rho 0.3696 query-only; 0.3772 content+query; 0.3675 LambdaMART | Query-source relevance carries more signal | Proprietary production-engine weighting |
IVRIS calculation based on the released modern anchor report: 0 of 3 interventions had a positive pooled estimate, or 0%. Two of three intervals, 66.7%, crossed zero; one of three, 33.3%, was negative with a 95% interval excluding zero. Formula: count ÷ 3 interventions × 100.
As of September 10, Google results for these queries describe topical authority, answer-first formatting, schema, statistics, citations, and other signals as GEO “ranking factors”; the GEO score AI citations SERP is led by commercial scoring tools. The preprint does not invalidate those ideas wholesale. It shows why a checklist needs current outcome validation before being treated as a citation predictor.
Why 0.114 Is Weak, Not Zero
Spearman correlation measures rank agreement. For each query, five candidate sources are ranked by score and by observed citation visibility. A value of 0.114 means there was a small positive tendency for those orderings to agree. It does not mean the score was 11.4% accurate.
Engine variation was wide: mean rho ranged from 0.0156 on GPT-OSS-120B to 0.1544 on Llama 3.3 70B. One cross-engine score should not be read as a universal rule.
The leakage correction is equally important. The first fitted evaluation allowed the same query to appear in training under one engine and testing under another. Query-disjoint folds reduced the full-model correlation from 0.5724 to 0.3625. IVRIS calculation based on the released strict-ranking artifacts: (0.5724 − 0.3625) ÷ 0.5724 = 36.7% relative reduction.
Corrected models retained more signal when the query was included, supporting a narrower conclusion: source selection is more query-dependent than a static page score captures. It also fits evidence that AI citation sets do not simply mirror Google’s organic top 10.
What B2B SEO and Content Teams Can Infer
Study finding: this query-blind score weakly predicted citation ordering, three historical GEO interventions produced no positive pooled effects, and query-conditioned predictors performed better.
Authors’ interpretation: page-only scores fit quality filtering better than citation prediction, and old causal anchors should be re-measured.
IVRIS analysis: treat a GEO score as a hypothesis generator unless it has current, query-disjoint outcome validation. “This page has desirable properties” and “this page will be cited for this query” are different claims.
Unknown: the study does not establish that GEO is ineffective across Google AI Overviews, ChatGPT Search, Perplexity, or every retrieval stack. It tests a controlled candidate-source setting, one score design, and three interventions; it does not measure traffic, conversions, pipeline, or revenue. The released package also uses stored engine snapshots, and it does not include the deterministic scorer needed to regenerate feature captures from raw page text.
Frequently Asked Questions
No. The study found a small positive relationship, not zero. The stronger conclusion is about scope: a query-blind page score was a weak predictor of which of five candidate sources received more citation visibility. Such a score may still be useful for quality control without being a reliable citation forecast.
No. The quotation and statistics estimates were slightly negative but their 95% confidence intervals included zero. Cite-sources was negative with an interval excluding zero. Those results apply to the authors’ paired, volume-controlled implementations and tested engine families, not every possible use of those content elements.
No. As of September 10, 2026, it is an arXiv preprint submitted September 7. The authors released reproducibility artifacts and a consistency audit, but those are not independent peer review or external replication. The findings should therefore be treated as preliminary evidence pending outside validation.
Not directly. The experiment measures answer-side visibility among supplied candidate sources. Production search adds retrieval, indexing, ranking, personalization, freshness, and proprietary behavior. This is evidence about the tested setup, not a universal ranking-factor study for every live AI search product.





