BlogStrategy
Strategy

Generative engine optimisation (GEO): what it is, what the research shows, and how to do it

Generative engine optimisation is the practice of improving content so AI systems — ChatGPT, Perplexity, Google AI Overviews, and Gemini — select it as a citation source. Here is the evidence base: what GEO tactics are confirmed, which are overhyped, and how to prioritise.

The Princeton KDD Lab published the first peer-reviewed study on generative engine optimisation in 2024. Since then, a body of controlled research has accumulated across nine independent studies. Some tactics are confirmed. Many are not. This is the evidence base for what actually works, scored by source quality.

2.4×
citation rate improvement from expert author attribution across all four AI platforms
Presence AI, 1,200 pages, 3,600 queries, ChatGPT + Perplexity + Google AI Overviews + Gemini, 90-day tracking.

What generative engine optimisation means

Generative engine optimisation (GEO) is the practice of improving content so AI systems select it as a cited source in their generated answers.

ChatGPT, Perplexity, Google AI Overviews, and Gemini are the four systems the term covers. Princeton's Knowledge Discovery and Data Mining lab coined and defined GEO in a peer-reviewed 2024 study that tested nine optimisation strategies across 9,679 queries from a 10,000-query dataset.

GEO is related to AEO (Answer Engine Optimisation), which predates the generative AI wave and covers optimising for featured snippets, voice search, and structured answer formats. GEO is the AI-specific evolution: optimising for retrieval-augmented generation (RAG) systems that synthesise answers from multiple sources and attribute citations.

What did the Princeton GEO study confirm?

Princeton's 2024 KDD study tested 9 content strategies against three AI answer engines and found only one, keyword stuffing, failed to help.

Citing relevant statistics increased visibility by 40% on average across the 3 systems Princeton tested; adding expert quotations added another 20%. The other six tested strategies, including citations, fluency improvements, and authoritative tone, produced smaller but positive gains, and combining several of them performed best of all 9.

Only keyword stuffing underperformed, confirming both that AI systems respond to content-quality signals in measurable ways, and that keyword-first optimisation (the foundation of traditional SEO) does not transfer to generative AI citation — a pattern Princeton's 2024 dataset of 9,679 queries reinforces.

The five most evidence-backed GEO tactics

1. Author credentials (Tier 1 by measured impact)

Author credentials produce the single largest measured impact on AI citation rates of any GEO tactic in this database.

Presence AI ran a 90-day controlled study tracking 1,200 pages and 3,600 queries across ChatGPT, Gemini, Perplexity, and Google AI Overviews.

Pages with expert authors and documented credentials achieved a 72% AI citation rate, versus 25% for pages with no author attribution — the gap Presence AI reported as a 2.4x difference, and the reason E-E-A-T author attribution ranks as the single highest-impact actionable GEO tactic in the database.

Implementation: every published page needs a visible author byline, a credentials statement naming specific expertise (not a vague bio), a link to an author page with documented background, and Person schema in the article markup, which matters specifically on Google AI Overviews.

2. Branded web mentions (Tier 1 by predictive correlation)

Branded web mentions predict AI Overview citation rates 3x more strongly than backlinks do, per Ahrefs.

Ahrefs analysed 75,000 brands and found branded mentions — references to a brand by name in third-party publications, with or without a hyperlink — carry a Spearman correlation of 0.664 with AI Overview citation rates, against 0.218 for backlinks.

The 0.664-vs-0.218 gap matters for where PR and content budget should go: link-building earns the weaker signal, while getting a brand mentioned by name across diverse third-party sources, even without a link back, earns the stronger one.

Implementation channels: digital PR, expert interviews, being cited as a source in industry publications, and brand mentions in aggregator or round-up articles. The goal is mentions across diverse sources, not link acquisition.

3. Content structure: heading structure and extractability

Heading structure is one of the strongest per-query predictors of AI citation rate in ChatGPT, per an 815,484-page AirOps dataset.

AirOps found that 7 to 20 subheadings, spanning 500 to 2,000 words, produced the highest consistent citation rate in ChatGPT across those 815,484 pages; both over-structured pages (too many headings) and under-structured prose (no internal navigation) underperformed that band.

Microsoft's Bing documentation corroborates this: strong descriptive headings are signals that help AI know where a complete idea starts and ends.

Treat each H2 as independently extractable: it should stand alone as a complete answer to the sub-question it addresses.

4. Specific, sourced claims with statistics

Sourced statistics were the single strongest lever Princeton's 2024 study measured, lifting AI visibility roughly 40% versus content without them.

Growth Memo's analysis of 21,482 ChatGPT citations found DATE and NUMBER are the two strongest positive entity signals in a page's first 1,000 characters: specific, verifiable claims create extraction points AI systems can cite with confidence, particularly early in a page, where retrieval concentrates.

The implementation is explicit: cite the source of every statistic, include the study's methodology and sample size where known, and name the specific percentage or count rather than approximating. Content written like a citable research finding (specific, attributed, falsifiable) outperforms content that reads like marketing copy.

5. Content freshness

Content freshness measurably affects AI citation, with Perplexity and ChatGPT both showing strong recency bias in what they cite.

Perplexity weights content published in the last 30 days most heavily of the four major AI platforms covered here. ChatGPT, when browsing is enabled, shows the same pattern: Amsive's analysis found 50% of AI-cited content is under 13 weeks old.

Content updated with new data or revised findings outperforms equivalent, older content on the same topic. For anyone maintaining an evergreen guide, this makes a periodic republish-with-new-data cadence a genuine citation lever for AI search, not just a freshness signal for Google's own ranking algorithm.

GEO tactics with weak evidence

Three tactics widely recommended for GEO are poorly supported by the research.

TacticRecommendation prevalenceEvidence qualityVerdict
FAQ schema markupVery highAccuraCast: 1.8% of cited pagesNot a meaningful AI citation driver
Meta descriptionHighWritesonic: 0/6 crawler readabilityNot read by AI crawlers
Open Graph tagsHighWritesonic: 0/6 crawler readabilityNot read by AI crawlers
JSON-LD structured dataVery highWritesonic: 0/6 crawler readabilityIndirect benefit via Google only
Word count maximisationModerateMixed (rate vs volume trade-off)Context-dependent

Source: AccuraCast (9,000 citation sources), Writesonic (62 elements, 6 crawlers), AirOps (n=815,484)

The AirOps dataset (815,484 pages) and a Writesonic crawl testing 62 webpage elements across six AI crawlers both confirm that meta descriptions, Open Graph tags, and JSON-LD schema markup are not read by most AI crawlers directly.

An AccuraCast study of 9,000 AI-cited pages found FAQ schema present on just 1.8% of them, despite being the most commonly recommended structured data type for AEO.

How does GEO differ by AI platform?

GEO is not uniform across AI systems. Platform architecture determines which signals reach the model.

Google AI Overviews builds on Google Search infrastructure. Traditional SEO signals (backlinks, domain authority, structured data) transfer more strongly here than to any other AI platform. Person schema and Article schema are relevant for Google AI Overviews in a way they are not for ChatGPT or Perplexity.

ChatGPT (GPT-4o with browsing) shows the weakest correlation with backlinks and the strongest preference for branded web mentions and author credentials. It also requires explicit GPTBot permission in robots.txt, making AI crawlability a prerequisite for citation.

Perplexity weights content freshness more aggressively than the other three platforms. It also produces the highest source overlap with Google (15.2%) of the ChatGPT/Claude/Gemini group, suggesting Google ranking transfers to Perplexity better than to other platforms.

Gemini uses Google infrastructure and shows citation patterns closer to Google AI Overviews than to ChatGPT or Perplexity. Standard SEO and GEO signals both apply.

How do you measure GEO performance?

Measuring GEO performance requires platform-specific monitoring, because Google Search Console does not capture AI citation data from ChatGPT, Perplexity, or Gemini standalone.

The available approaches: manual query testing (ask target questions directly in each AI platform and record citation sources), AI answer monitoring platforms (track which sources are cited in AI responses to a monitored keyword set), and share-of-voice tracking adapted for AI search. As of mid-2026, no single tool covers all four major platforms with equivalent depth.

The minimum viable GEO measurement: monthly manual audits of your top 20 target queries in ChatGPT and Perplexity, plus Google AI Overviews, recording which pages are cited and which competitors appear. This gives a trend line without requiring specialised tooling.

Google-ranking tactics don't transfer to AI citation

Backlinks, keyword density, and FAQ schema drive Google ranking but are weak or irrelevant for AI citation rates — the source pools, the signals, and the measurement approach are all different disciplines wearing similar-looking names.

Author attribution, branded web mentions, heading structure, and specific sourced claims are what drive GEO performance instead. Each is well-evidenced across multiple independent studies and should be treated as the foundational layer of any content strategy aimed at AI search visibility.

Frequently asked questions

Is generative engine optimisation (GEO) the same as AEO?

GEO and AEO are related but distinct: AEO targets pre-generative formats like featured snippets and voice search, while GEO targets sources an AI cites inside a fully generated answer. Most tactics that help one also help the other, but optimising only for snippet capture won't automatically earn AI citations.

How is GEO different from SEO?

SEO targets Google's ranking algorithm (PageRank, domain authority, backlinks, keyword density). GEO targets AI retrieval mechanisms that select sources for synthesised answers. The core difference: a peer-reviewed arXiv study found GPT-4o's source pool overlaps with Google by only 4%. What predicts Google ranking (backlinks) barely predicts AI citation (0.218 Spearman correlation), while branded web mentions, weaker for Google, have 3x stronger correlation with AI citation rates (0.664). The tactics and the target systems are different.

Does author-credential attribution need Person schema to work?

Person schema is not a universal requirement for author-credential attribution to work — treat it as the last step, not the first. A visible byline, a credentials statement, and a linked author page do the work across every platform in the underlying study, while Person schema markup only adds incremental value specifically on Google AI Overviews. If you have to choose where to spend limited implementation time, the visible on-page elements matter more broadly than the structured-data wrapper around them.

Does ChatGPT require special crawler permission before it can cite your content?

ChatGPT (GPT-4o with browsing) requires explicit GPTBot permission in robots.txt before it can crawl a page at all — check your robots.txt for a GPTBot disallow rule now, since the fix takes minutes but unblocks everything else. Without that permission, tactics like author attribution and branded mentions never get the chance to matter, because the page is never retrieved in the first place. It is a one-time technical fix, not an ongoing content practice, which makes it easy to overlook.

What are the most common GEO mistakes?

FAQ schema is the most common GEO mistake — it appears in only 1.8% of AI-cited pages despite being the most widely recommended structured data type for AEO. Publishing anonymous content is the second: unattributed pages achieve 2.4x fewer citations than pages with expert author attribution. The third is applying Google SEO metrics to measure GEO performance, since Google Search Console data does not capture ChatGPT or Perplexity citation rates — that requires separate monitoring with AI answer tracking tools.

BE

BetterAISearch Editorial Team

BetterAISearch

The BetterAISearch team synthesises peer-reviewed studies, platform documentation, and independent research into actionable, scored tactics.

Get the next finding first.

New research moves scores. We send what changed and what it means, before it becomes LinkedIn noise.