Generative Engine Optimization: What the Original GEO Study Measured, What Google Says, and the Position Arithmetic in Between
· Royking Niba
Generative engine optimization, or GEO, is the practice of making your content more likely to be used, quoted and cited inside answers written by AI search systems such as Google’s AI Overviews and AI Mode, ChatGPT search and Perplexity. The term comes from a research paper first published in November 2023 and accepted to KDD 2024, which tested nine ways of rewriting a web page and measured how much of each AI answer came from it. The three edits that helped most were adding quotations, adding statistics and citing sources. Keyword stuffing made things worse. Google, meanwhile, says there are no special optimizations needed for its own AI features. Both statements are true, and this page explains how they fit together, with a worked example of the visibility metric the study used.
Where the term comes from
The paper is “GEO: Generative Engine Optimization” by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, first posted to arXiv on 16 November 2023 and revised to its current version on 28 June 2024. Its abstract describes GEO as “the first novel paradigm to aid content creators” in improving their visibility in generative engine responses, and reports strategies that boost visibility by up to 40 percent, with results that vary by domain.
How it was tested matters more than the headline. The authors built a benchmark, GEO-bench, of 10,000 queries across 25 domains, split 8,000 for training, 1,000 for validation and 1,000 for testing. For each query they took the top five Google results, had a language model (gpt-3.5-turbo) write an answer citing them, then rewrote one source using each strategy and measured how its share of the answer changed. They also repeated the test on Perplexity, a live AI search engine, and report visibility improvements of up to 37 percent there.
What the study found, method by method
The headline figures from the paper’s main results, using its position-adjusted word count metric. The baseline is an unmodified source. The absolute change in percentage points is my arithmetic from the paper’s figures.
| Rewrite strategy | Visibility score | Change in points | Relative change |
|---|---|---|---|
| No change (baseline) | 19.3 | 0.0 | 0% |
| Quotation addition | 27.2 | +7.9 | +41% |
| Statistics addition | 25.2 | +5.9 | +31% |
| Fluency optimization | 24.7 | +5.4 | +28% |
| Cite sources | 24.6 | +5.3 | +27% |
| Keyword stuffing | 17.7 | -1.6 | -8% |
The other four strategies tested were making the text more authoritative in tone, easier to understand, richer in unique words, and richer in technical terms. The pattern across all nine is the useful part. The edits that worked add evidence: a quotation from someone credible, a number, a named source. The edit that failed adds repetition. That is the same line Google’s spam policies draw for classic search, which is why I do not treat GEO as a separate discipline so much as a stricter grader of the same qualities.
What Google says about its own AI features
Quoted from Google Search Central’s page on AI features and your website, last updated 10 December 2025, checked today:
There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.
Google Search Central, AI features and your website
The same page sets one hard eligibility rule: “To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet”. It explains that both features may use a “query fan-out” technique, issuing several related searches across subtopics and sources to build one answer. And it says traffic from these features is reported inside the Search Console Performance report under the Web search type, not separately, adding: “We’ve seen that when people click from search results pages with AI Overviews, these clicks are higher quality (meaning, users are more likely to spend more time on the site).”
Put the two sources side by side and they do not conflict. Google is saying there is no special markup, file or tag that unlocks its AI features: a page has to be indexable with a snippet, and then it competes on usefulness. The GEO study is saying that, among pages already in the candidate set, the ones carrying quotable evidence get used more. Neither says you can rank in an AI answer from a page that cannot rank in search. That is also why an llms.txt file is not a GEO strategy.
The position arithmetic: where you are cited matters
The study’s main metric is not “were you cited”. It is a position-adjusted word count: for every sentence in the answer that cites your source, take its length and discount it by how late it appears, using a factor of e to the power of minus (position divided by the number of sentences). Then divide by the total words in the answer. Earlier sentences count for more.
Here is a worked example. The answer and its numbers are invented to show the mechanics; the formula is the paper’s. Assume an AI answer of 10 sentences, 20 words each, 200 words in all, with sentences numbered 1 to 10.
| Your source is cited in | Discount applied | Visibility share |
|---|---|---|
| Sentence 1 only | 0.905 | 9.0% |
| Sentence 5 only | 0.607 | 6.1% |
| Sentence 10 only | 0.368 | 3.7% |
| Sentences 9 and 10 | 0.407 and 0.368 | 7.7% |
| Sentences 1 and 2 | 0.905 and 0.819 | 17.2% |
| Sentences 1, 2 and 3 | 0.905, 0.819 and 0.741 | 24.6% |
Two things fall out of this. First, the opening sentence is worth 2.46 times the closing one, so being the source of the direct answer is worth far more than being a supporting footnote. Second, being cited twice at the end (7.7 percent) scores less than being cited once at the start (9.0 percent). The practical reading: the content that gets used first is the content that answers the question in its first lines, in a form an AI system can lift as a sentence. That is why I open pages with a bolded, self-contained answer and put the evidence (the quote, the number, the source) inside that opening rather than further down.
The model’s limits: real answers vary in sentence length, AI Overviews and AI Mode do not publish how they order sources, and the study’s metric is a research measure, not something you can read off Search Console. Use it to understand the shape of the problem, not as a forecast.
How I apply GEO on a real site
| Step | What I check | Why |
|---|---|---|
| 1. Eligibility | Is the page indexed and shown with a snippet? | Google’s stated requirement for a supporting link |
| 2. Answer first | Does the first paragraph answer the query on its own? | Early sentences carry the most weight in the visibility metric |
| 3. Evidence density | Is there a quote, a figure or a named source in the answer itself? | The three strongest strategies in the study all add evidence |
| 4. No stuffing | Is the target phrase repeated for its own sake? | The only strategy in the study that lowered visibility |
| 5. Fan-out coverage | Does the page answer the obvious follow-up questions? | Google says its AI features issue related searches across subtopics |
| 6. Measurement | Are Web search clicks and impressions moving on the pages you edited? | AI feature traffic is folded into the Performance report |
Steps three and four are where GEO and penalty recovery meet. Google’s spam policies target tactics built on repetition: templated pages, doorway pages, keyword-stuffed copy. The study’s results suggest AI answers punish the same habits by simply not using that text. My Google penalty recovery guide covers the search side of that cleanup.
For the closely related ideas, see answer engine optimization, which is the same answer-first discipline applied to featured snippets and voice, how to rank in AI Overviews, and how to get cited by ChatGPT.
Leave a Reply