How do AI systems decide which companies to recommend?
What is documented by the providers, what is strong inference, and what nobody outside those companies actually knows.
No AI provider publishes how its assistant selects or ranks companies in an answer, and there is no equivalent of a search ranking-factors document. Anyone stating the mechanism with confidence is inferring.
What can be said accurately is that answers are produced from two sources: what the model absorbed during training, and what it retrieves live at question time. Both are shaped by how widely and how consistently a claim about a company appears across sources those systems read — which is why corroboration, rather than assertion, is the lever that appears to matter most.
Documented: how content reaches these systems
The retrieval pathway is well documented, because the providers publish crawler documentation in order to give site owners control. OpenAI operates OAI-SearchBot for surfacing sites in ChatGPT search, ChatGPT-User for user-initiated fetches, and GPTBot for model development. Anthropic operates Claude-SearchBot for search quality, Claude-User for user-initiated retrieval, and ClaudeBot for training. Perplexity operates PerplexityBot for indexing and Perplexity-User for user-triggered fetches — and states that the latter generally ignores robots.txt. Google uses Googlebot for Search, with Google-Extended as a separate control for Gemini training that, by Google’s own statement, does not affect inclusion or ranking in Search.
That is the extent of what is documented. It tells you how to be readable. It says nothing about how a company is chosen once several are readable.
Strong inference: what the observable behaviour implies
Four inferences are well supported by consistent observation and by the published research, and we label them as inference rather than fact.
Consistency across independent sources appears to dominate. Companies described the same way across many independent sources are described that way by models, including in sessions with no live retrieval. This is the most robust pattern we observe and the least surprising: it is what "learning from a corpus" means.
Specificity aids retrieval and citation. The GEO research found that adding citations, quotations and statistics improved citation rates, while keyword-oriented edits did not. Content that makes a specific, attributable claim is easier for a system to use.
Category framing shapes the comparison set. The competitors a model names alongside you reveal the category it has filed you under. That filing is frequently wrong, and it is changeable — but only by changing what independent sources say you are, not by changing your own tagline.
Recency matters differently by system. Assistants that retrieve live reflect new pages quickly; answers drawn from model priors lag by months. The same company can be described in two different ways by two systems on the same day, and both are "correct" given their inputs.
Unknown: what nobody outside these companies can tell you
Whether there is an explicit ranking function for entities, and what it weights. Whether brand mentions without links carry weight, and how much. How much any individual source is trusted relative to another. Whether commercial arrangements influence which companies appear — beyond disclosed advertising products, none of which are known to affect organic recommendation. And whether any of the above will remain true next quarter.
It is worth being blunt about this, because the market is full of confident mechanistic claims presented as inside knowledge. There is no equivalent of the Google ranking-factors literature for assistants, and the systems change without release notes.
You cannot optimise against a specification that does not exist. What you can do is make one specific, true claim easy to find, easy to verify, and repeated by sources that are not you — and then measure whether the answers change. That is robust to the mechanism being unknown, which is the point.
Sources
- OpenAI — crawler documentation (GPTBot, OAI-SearchBot, ChatGPT-User, OAI-AdsBot)
- Anthropic — “Does Anthropic crawl data from the web, and how can site owners block the crawler?”
- Perplexity — crawler documentation (PerplexityBot, Perplexity-User)
- Google Search Central — Google crawlers, fetchers and user-triggered fetchers
- Aggarwal et al. — “GEO: Generative Engine Optimization” (KDD 2024)
- Vercel — “The rise of the AI crawler” (December 2024)
Related questions.
The questions that usually come next, answered to the same standard.
Keep reading.
This page answers the general question.
The specific one — what these systems say about your company, and where those words came from — takes a working session and about a week. The finding is yours whether or not we go further.
The first call is complimentary — and the finding is yours to keep.