GEO Score Shows Weak 0.11 Correlation With AI Citations

Home News GEO Score Shows Weak 0.11 Correlation With AI Citations
Digital Marketing

A September 7, 2026 arXiv preprint by Elisha Bajemon and Andre-Louis Rochet tested a question GEO scorecards often skip: does a page-only content score predict which source a modern generative engine will cite? In the original seven-engine outcome set, the pooled within-query Spearman correlation was only 0.114 across 777 evaluable query-engine groups. The paper also ... <a title="GEO Score Shows Weak 0.11 Correlation With AI Citations" class="read-more" href="https://ivristech.com/geo-score-ai-citations-correlation/" aria-label="Read more about GEO Score Shows Weak 0.11 Correlation With AI Citations">Read more</a>

PK
September 11, 2026 5 min

A September 7, 2026 arXiv preprint by Elisha Bajemon and Andre-Louis Rochet tested a question GEO scorecards often skip: does a page-only content score predict which source a modern generative engine will cite? In the original seven-engine outcome set, the pooled within-query Spearman correlation was only 0.114 across 777 evaluable query-engine groups.

The paper also re-tested three original GEO interventions: quotations, statistics, and source citations. None produced a positive pooled effect. Two confidence intervals included zero; cite-sources was negative with an interval excluding zero.

This does not mean “GEO does not work.” This query-blind score weakly predicted citation ordering, while the three historical interventions did not transfer positively. Query-conditioned models reached about 0.37 to 0.38.

Direct answer — Do GEO scores predict AI citations?

In this preprint, a deterministic content-only GEO score had a weak positive relationship with citation visibility: within-query Spearman rho was 0.114 on seven original engine arms and 0.118 in a three-model GPT-5.x replication. That is not zero, and it is not 11.4% accuracy. It means the score only weakly preserved the engines’ ordering of five candidate sources for the same query.

Key Takeaways

  • The work is an arXiv preprint, not peer reviewed.
  • The fixed score reached rho = 0.114 across 777 evaluable groups; three GPT-5.x arms replicated at 0.118 across 450 groups.
  • Quotation, statistics, and cite-sources edits produced 0 of 3 positive pooled effects.
  • The ranking data contain 229 unique queries and 1,269 query-engine groups, with five candidate sources per group.
  • Correcting query leakage cut a fitted-model result from 0.5724 to 0.3625.

What the Study Actually Tested

The study separates two objects. First is a deterministic 0–100, query-blind score built from 11 text features: entropy, information and entity density, semantic coherence, self-containment, citation F1, statistic density, MMR, nDCG, redundancy, and quotable density. Second is a query-conditioned predictor that adds query-source relevance and asks how well it can reproduce the engine’s ordering of five candidate sources.

The denominators differ. The 500-source benchmark tests score-gaming resistance, not rho 0.114. Citation-ranking data contain 229 unique queries and 1,269 query-engine groups. The original seven-arm analysis has 777 evaluable groups; GPT-5.5, GPT-5.4, and GPT-5.4-mini add 150 each.

The outcome is a PAWC-like visibility share in generated answers, not Google rank, traffic, leads, or revenue. Keep those layers separate in AI search visibility measurement.

GEO claim / tacticWhat was actually testedObserved resultEvidence supportsStill unknown
A page-only GEO score predicts citationsQuery-blind score versus within-query source visibilityrho 0.114, n=777; GPT-5.x rho 0.118, n=450Weak positive ordering relationshipWhether commercial GEO scores predict live-search citations
Add quotationsPaired, volume-controlled quotation edit-0.325 pp; 95% CI -0.841 to 0.171; n=1,531No positive pooled effect detectedOther styles, domains, prompts, or retrieval systems
Add statisticsPaired, volume-controlled statistics edit-0.276 pp; 95% CI -0.970 to 0.425; n=1,087No positive pooled effect detectedWhether original or highly query-relevant data differ
Add source citationsPaired, volume-controlled cite-sources edit-0.793 pp; 95% CI -1.533 to -0.138; n=1,087No positive effect; pooled estimate was negativeWhether source quality or implementation changes the result
Query context mattersQuery-aware models on query-disjoint foldsrho 0.3696 query-only; 0.3772 content+query; 0.3675 LambdaMARTQuery-source relevance carries more signalProprietary production-engine weighting

IVRIS calculation based on the released modern anchor report: 0 of 3 interventions had a positive pooled estimate, or 0%. Two of three intervals, 66.7%, crossed zero; one of three, 33.3%, was negative with a 95% interval excluding zero. Formula: count ÷ 3 interventions × 100.

As of September 10, Google results for these queries describe topical authority, answer-first formatting, schema, statistics, citations, and other signals as GEO “ranking factors”; the GEO score AI citations SERP is led by commercial scoring tools. The preprint does not invalidate those ideas wholesale. It shows why a checklist needs current outcome validation before being treated as a citation predictor.

Why 0.114 Is Weak, Not Zero

Spearman correlation measures rank agreement. For each query, five candidate sources are ranked by score and by observed citation visibility. A value of 0.114 means there was a small positive tendency for those orderings to agree. It does not mean the score was 11.4% accurate.

Engine variation was wide: mean rho ranged from 0.0156 on GPT-OSS-120B to 0.1544 on Llama 3.3 70B. One cross-engine score should not be read as a universal rule.

The leakage correction is equally important. The first fitted evaluation allowed the same query to appear in training under one engine and testing under another. Query-disjoint folds reduced the full-model correlation from 0.5724 to 0.3625. IVRIS calculation based on the released strict-ranking artifacts: (0.5724 − 0.3625) ÷ 0.5724 = 36.7% relative reduction.

Corrected models retained more signal when the query was included, supporting a narrower conclusion: source selection is more query-dependent than a static page score captures. It also fits evidence that AI citation sets do not simply mirror Google’s organic top 10.

What B2B SEO and Content Teams Can Infer

Study finding: this query-blind score weakly predicted citation ordering, three historical GEO interventions produced no positive pooled effects, and query-conditioned predictors performed better.

Authors’ interpretation: page-only scores fit quality filtering better than citation prediction, and old causal anchors should be re-measured.

IVRIS analysis: treat a GEO score as a hypothesis generator unless it has current, query-disjoint outcome validation. “This page has desirable properties” and “this page will be cited for this query” are different claims.

Unknown: the study does not establish that GEO is ineffective across Google AI Overviews, ChatGPT Search, Perplexity, or every retrieval stack. It tests a controlled candidate-source setting, one score design, and three interventions; it does not measure traffic, conversions, pipeline, or revenue. The released package also uses stored engine snapshots, and it does not include the deterministic scorer needed to regenerate feature captures from raw page text.

Frequently Asked Questions

No. The study found a small positive relationship, not zero. The stronger conclusion is about scope: a query-blind page score was a weak predictor of which of five candidate sources received more citation visibility. Such a score may still be useful for quality control without being a reliable citation forecast.

No. The quotation and statistics estimates were slightly negative but their 95% confidence intervals included zero. Cite-sources was negative with an interval excluding zero. Those results apply to the authors’ paired, volume-controlled implementations and tested engine families, not every possible use of those content elements.

No. As of September 10, 2026, it is an arXiv preprint submitted September 7. The authors released reproducibility artifacts and a consistency audit, but those are not independent peer review or external replication. The findings should therefore be treated as preliminary evidence pending outside validation.

Not directly. The experiment measures answer-side visibility among supplied candidate sources. Production search adds retrieval, indexing, ranking, personalization, freshness, and proprietary behavior. This is evidence about the tested setup, not a universal ranking-factor study for every live AI search product.

Share
PK
Written by
Priyanshi Kharwade
Priyanshi Kharwade — B2B News & Content | Ivris Tech
Content writer covering B2B news and market trends. Communication student with a background in digital marketing and editorial writing. Tracks the developments that matter for B2B operators.

Get B2B marketing insights weekly

Strategies, frameworks, and tools — no fluff. Join operators who read Ivris Tech.

No spam. Unsubscribe anytime.
Link copied!