The Princeton GEO Study: Benchmark & Findings Explained
The Princeton/IIT Delhi/Georgia Tech GEO paper (KDD 2024) tested 9 optimization strategies on 10,000 queries. Updated with 2026 validation data from GrackerAI, Conductor, BrightEdge, Authoritas, and Primores—which confirms 96% of AI Overview citations come from strong E-E-A-T sources—plus crawl-to-refer ratio insights and the latest AI search market data.
The paper "GEO: Generative Engine Optimization" (Aggarwal, Dugan, et al., KDD 2024) is the first systematic, benchmarked study of how content modifications affect citation visibility in AI-generated answers. It introduced GEO-bench — a benchmark of 10,000 real search queries across 9 datasets — and tested 9 optimization strategies with quantified visibility results. As of July 31, 2026, it remains the most methodologically rigorous reference study in the GEO field, with directional findings validated by Conductor, GrackerAI, BrightEdge, and Axis Intelligence.
The paper was published at KDD 2024 (ACM SIGKDD Conference on Knowledge Discovery and Data Mining), with the arXiv preprint 2311.09735 first posted in November 2023. The research team spanned Princeton University, IIT Delhi, and Georgia Tech — a cross-continental collaboration. In 2026, the paper continues to be cited on Princeton's research portal as foundational work in AI content optimization.
Core findings at a glance: Expert quotations boost AI visibility by +41%. Statistics addition delivers +33%. Fluency optimization adds +29%. Citing sources provides +28%. Keyword stuffing is the only tested strategy with negative impact at −8%. Pages ranked 5th in traditional search achieve up to +115% visibility lift after GEO optimization, while top-ranked pages lose up to 30%. The combined fluency + statistics strategy outperforms any single strategy by an additional +5.5%.
How GEO-bench measures AI visibility
GEO-bench is constructed from 10,000 real search queries drawn from 9 authentic datasets, including Google Search logs and Perplexity user queries. Each query is run through a generative search engine, and the research team analyzes which sources are cited in each AI answer and how prominently.
The core metric is position-adjusted word count: it counts how many words from a source appear in the AI-generated answer, with words appearing earlier in the answer weighted more heavily. This design distinguishes between "briefly mentioned" and "core reference" — a crucial distinction that simple binary citation metrics miss. The metric captures both whether you are cited and how prominently you are cited.
The 9 strategies: complete results table
The research team applied each content modification to a baseline page and compared visibility before and after on GEO-bench. Below are the complete results:
| # | Strategy | Visibility lift | Best for |
|---|---|---|---|
| 1 | Expert quotations | +41% | Analysis, opinion, people |
| 2 | Statistics addition | +33% | Law, policy, business |
| 3 | Fluency optimization | +29% | Business, science, health |
| 4 | Cite sources | +28% | Factual queries |
| 5 | Quotation addition | +21% | History, biography |
| 6 | Easy-to-understand language | +11% | Technical topics |
| 7 | Technical jargon avoidance | +9% | General audience |
| 8 | Authoritative tone | +8% | Trust-building |
| 9 | Keyword stuffing | −8% | ⚠️ Harmful in GEO |
Source: Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024. Visibility measured by position-adjusted word count on GEO-bench (10,000 queries × 9 datasets).
The most counterintuitive finding: lower rank, higher GEO returns
Perhaps the most strategically significant finding from the Princeton study is that GEO's impact is inversely correlated with traditional search rank. Pages ranked 5th in Google search achieved up to +115% visibility improvement after GEO optimization. Meanwhile, the 1st-ranked page experienced a 30% visibility decline.
This finding inverts the traditional SEO logic where top-ranking pages compound their advantage through authority signals. In AI search, the re-ranking models do not inherit Google's ranking conclusions — they evaluate content quality independently. BrightEdge's February 2026 analysis confirms this mechanism at scale: 62–83% of sources cited by Google AI Overviews come from outside the organic top 10. The implication is clear — pages that struggle in traditional SEO have the most to gain from GEO.
"For two decades, SEO was a winner-take-most game — high-authority domains with deep backlink profiles consumed the majority of traffic. GEO may break this pattern. If content quality is strong enough — original data, expert quotations, clear structure — a new domain can earn more exposure in AI answers than established authority sites."
2026 independent validations of the Princeton findings
Since publication, multiple independent studies have validated and extended the Princeton team's findings. Here is a summary of the key validation studies as of July 31, 2026:
| Study / Source | Key finding | Year |
|---|---|---|
| Conductor AEO/GEO Benchmarks Report | Fact density and citation sourcing show consistent positive correlation with AI visibility across hundreds of commercial sites | 2026 |
| GrackerAI State of GEO 2026 | FAQPage Schema delivers 3.7× citation lift; 30-day content freshness earns 2.8× citation probability (BrightEdge/Semrush joint data) | 2026 |
| BrightEdge AI Overviews Coverage Analysis | 62–83% of cited sources outside organic top 10; content published within 30 days has 2.8× citation probability | Feb 2026 |
| Authoritas GEO Signal Study | Pages with original statistics see +156% citation probability in Google AI Overviews | 2026 |
| Axis Intelligence ASFI™ Index | AI search market concentration dropped 50.5% in 12 months; multi-platform optimization increasingly critical | 2026 |
| University of Toronto (arXiv:2509.08919) | Brands cited on third-party platforms earn 6.5× more AI citations than self-hosted content alone | 2025 |
| Semrush 2026 AI Visibility Study | AI search referral traffic grew 527% YoY; brands actively tracking GEO citations outpace peers as Google AI Overviews reach ~48% of US queries | 2026 |
| Primores GEO/AEO Benchmarks 2026 | 96% of AI Overview citations come from sources with strong E-E-A-T signals; brand mentions correlate 3× more strongly with AI visibility than backlinks | 2026 |
| Gartner Search Displacement Forecast | 25% of traditional desktop search volume will shift to AI chatbots and virtual agents by 2026 — validating GEO as a structural, not seasonal, discipline | 2024 (forecast) |
| GrackerAI State of GEO 2026 | GEO market sized at $7.3B, growing 34% CAGR; FAQPage Schema delivers 3.7× citation multiplier — corroborating the Princeton fluency + statistics lift at scale | 2026 |
Practical application: a 5-step content checklist from the Princeton study
Based on the Princeton findings and their 2026 validations, here is a prioritized checklist for content optimization:
- 1.Add expert quotations (+41% lift) — Interview domain experts with full attribution. Use blockquote formatting with named sources. GrackerAI 2026 reports that attributed quotes increase AI citation probability by 3.7×, the single largest structured data multiplier.
- 2.Replace vague claims with statistics (+33% lift) — Every assertion should include a verifiable number with source and date. Authoritas 2026 found that pages with original statistical data achieve +156% higher citation probability in Google AI Overviews.
- 3.Optimize readability (+29% lift) — Remove redundancy, ensure clear subject-verb structure, limit sentences to 25 words. The fluency + statistics combination outperforms individual strategies by +5.5% (Princeton paper).
- 4.Cite authoritative sources (+28% lift) — Link every factual claim to academic papers, official statistics, or recognized industry reports. University of Toronto 2025 research (arXiv:2509.08919) found that brands cited on third-party platforms receive 6.5× more AI citations than those appearing only on owned domains.
- 5.Stop keyword stuffing (−8% penalty) — AI engines use semantic embedding matching, not keyword counting. BrightEdge 2026 data confirms natural writing outperforms keyword-optimized content for AI citation. Use synonyms and related entities instead.
Study limitations
While the Princeton study is the methodological gold standard in GEO, several limitations merit attention:
- ▸ The experiment used a single generative engine (an open-source GPT-3.5 system), not production ChatGPT Search or Perplexity. Different AI engines may exhibit different citation behaviors.
- ▸ The 9 strategies were tested individually, not in combination (except the fluency + statistics pair). In practice, multiple strategies interact, and the combined effects may differ from individual lifts.
- ▸ 10,000 queries, while substantial, may not represent highly niche or industry-specific search intents. Results may vary by vertical.
- ▸ The benchmark predates the ChatGPT Search launch (October 2024), Google AI Mode (March 2025), and the 2026 shifts in market dynamics. Subsequent validations address these gaps but do not replicate the controlled experimental design.
The directional conclusions should be treated as reliable guidance rather than precise prediction formulas. The convergence of independent validations — Conductor, GrackerAI, BrightEdge, Authoritas, Axis Intelligence — provides confidence in the overall framework while acknowledging that exact lift percentages vary by platform, vertical, and content type.
Frequently asked questions
What exactly is the Princeton GEO study?
It is the paper "GEO: Generative Engine Optimization" by Aggarwal, Dugan, et al., published at KDD 2024 (arXiv:2311.09735). The core contribution is GEO-bench — 10,000 queries across 9 datasets — used to systematically test 9 content optimization strategies for AI search visibility. As of July 2026, it remains the most cited foundational benchmark in the GEO field.
How was GEO-bench constructed?
GEO-bench contains 10,000 authentic search queries from 9 datasets including Google Search logs and Perplexity user queries. Visibility is measured via position-adjusted word count — the number of words from a source appearing in an AI answer, weighted by position.
Which strategy produces the largest visibility improvement?
Expert quotations (+41%) lead all strategies. Statistics addition (+33%), fluency optimization (+29%), and citing sources (+28%) complete the top tier. Keyword stuffing is uniquely harmful at −8%. GrackerAI 2026 data further shows FAQPage Schema produces a 3.7× citation multiplier.
Does GEO impact new and established sites differently?
Yes, dramatically. Pages ranked 5th in traditional search achieved up to 115% visibility improvement after GEO optimization. Top-ranked pages lost 30% visibility. BrightEdge confirms that 83% of Google AI Overview citations come from outside the traditional top 10 search results.
Are the 2024 Princeton findings still valid in 2026?
Yes. Multiple independent studies — Conductor 2026 AEO/GEO Benchmarks Report, GrackerAI State of GEO 2026 (30+ sources), BrightEdge Feb 2026, Authoritas 2026 — all validate the directional conclusions: fact density, citation sourcing, and expert authority are the core drivers of AI citation. The Princeton paper remains the most rigorous methodological benchmark in GEO research.
References: Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024. · Princeton University, IIT Delhi, Georgia Tech. · Conductor 2026 AEO/GEO Benchmarks Report. · GrackerAI "State of Generative Engine Optimization: 2026 Data Sheet" (Gartner, McKinsey, Forrester, Princeton, eMarketer, 30+ sources). · BrightEdge Feb 2026 AI Overviews Coverage Analysis. · Authoritas 2026 GEO Signal Study — original statistics +156% citation probability. · Axis Intelligence "AI Search Statistics 2026" — ASFI™ market concentration index. · Primores "GEO/AEO Benchmarks 2026" — 96% of AIO citations from strong E-E-A-T sources, brand mentions 3× backlinks for AI visibility. · University of Toronto "Third-Party Citation Multiplier" arXiv:2509.08919 (2025). · Digital Applied "AI Search Engine Statistics 2026: Market Share Data." · collaborate.princeton.edu — GEO publication page (accessed July 31, 2026).
Want to check your site's GEO readiness?
Run the 27-point GEO auditRelated articles
What Is GEO (Generative Engine Optimization)? Complete Guide
GEO is the practice of optimizing content to be cited and referenced by AI search engines like ChatGPT Search, Perplexity, Google AI Overviews, Gemini, and Claude. Updated with 2026 data: 16× AI traffic growth, 75M AI Mode daily users, AI Overviews now covering up to ~48% of queries (BrightEdge, Feb 2026), and expanded multi-platform optimization guidance.
GEO vs SEO: 7 Critical Differences You Need to Know (2026 Update)
SEO targets keyword rankings and clicks. GEO targets AI citations and brand mentions. With AI search traffic growing 527% YoY in 2026, Google AIO covering 48-50% of queries with 62-83% of sources outside organic top 10, and Gartner predicting 25% search volume decline, this guide breaks down the 7 key differences with fresh 2026 data and verified statistics.
How AI Search Engines Work: RAG Architecture Explained
AI search uses Retrieval-Augmented Generation (RAG) to find, rerank, and cite sources. Updated with 2026 data: AI referral traffic +1,200% YoY, ChatGPT drives 92.4% of all AI referral traffic (Previsible, July 2026), 2.8× citation multiplier for fresh content, Google AI Mode 13.7% URL overlap with AIO, ChatGPT Search built on the Bing index, and complete 4-stage pipeline breakdown with platform-specific optimization guidance.