What is generative engine optimization (GEO)?
The term, where it came from, what the evidence supports, and the part of it that is mostly folklore.
Generative engine optimization (GEO) is the practice of making a company, product or page more likely to be surfaced, described accurately, and cited by AI systems that answer questions directly — ChatGPT, Claude, Gemini, Perplexity and Google’s AI Overviews — rather than returning a list of links.
The term was introduced in a 2023 academic paper by Aggarwal and colleagues, later published at KDD 2024, which tested content modifications against a generative search engine and measured changes in citation visibility. In commercial use it has since expanded to cover almost any work aimed at AI answer visibility, including a good deal that the original research does not support.
What the original research actually found
The GEO paper tested a set of content edits against a generative engine and measured how often a source was cited in the resulting answer. The interventions that helped most were the unglamorous ones: adding citations to credible sources, adding quotations, and adding statistics. Interventions aimed at keyword density — the reflex of classical SEO — performed poorly or not at all.
That result is worth sitting with, because it is the opposite of how GEO is usually sold. The paper’s finding is that substance signals move citation, not surface optimisation. Content that quotes, cites and quantifies is easier for a generative system to use, because it can be lifted into an answer with its provenance attached.
Two caveats belong with that finding. The study was conducted on a specific generative engine at a specific point in time, and these systems change without notice. And a controlled improvement in citation rate on a benchmark is not the same as being recommended to a buyer in a live commercial market.
What GEO covers in practice
In commercial use the term now spans four fairly different activities, and conflating them is the source of most of the confusion in the market.
Legibility. Making sure a machine can read the page at all: content present in the served HTML, crawler access permitted, structure and entity markup coherent. This is a precondition rather than a strategy, and it is the one part that is unambiguously technical.
Substance. Making the page worth quoting: specific claims, real numbers, cited sources, direct answers stated early rather than buried under an introduction. This is what the research supports most directly.
Corroboration. Getting the claim repeated by independent sources the systems already read. This is the slowest work and, in our view, the one that most determines the outcome — and it is barely technical at all.
Measurement. Tracking what the systems say over time. Necessary, frequently oversold, and prone to mistaking session-to-session variance for progress.
What is mostly folklore
A number of widely repeated tactics have little or no evidence behind them. The clearest example is llms.txt, a proposed file for giving language models a curated map of a site. Ahrefs studied 137,210 domains and found that 97% of llms.txt files received no requests at all, and that AI bots never went looking for the file on sites that did not have one. It costs almost nothing to publish, and it does almost nothing.
A second is the belief that adding FAQ schema, or any schema, causes citation. Structured data helps a system parse a page and can affect eligibility for certain search features. There is no published evidence that it causes a generative system to recommend a company, and it plainly does not compensate for having nothing distinctive to say.
A third is guaranteed placement. No model provider sells access to the ranking in an assistant’s answer. Any firm offering a guarantee is describing an outcome it does not control.
Our position on the term
We use "GEO" descriptively because buyers search for it, and we do not organise our work around it. The label frames the problem as an optimisation exercise adjacent to SEO, which leads teams to spend on structure and content volume when the binding constraint is almost always what they are claiming and who else repeats it.
The more accurate framing is that AI systems are a new audience with unusual reading habits: they never ask a follow-up question, they discount self-description, and they reward claims they can verify elsewhere. That is a positioning problem with a technical floor — which is why we call the work machine-readable positioning rather than optimisation.
Related questions.
The questions that usually come next, answered to the same standard.
Keep reading.
This page answers the general question.
The specific one — what these systems say about your company, and where those words came from — takes a working session and about a week. The finding is yours whether or not we go further.
The first call is complimentary — and the finding is yours to keep.