Short answer
AI assistants name a business when it is described consistently across sources they trust — third-party pages more than your own. To get cited: allow the retrieval bots (OAI-SearchBot, Claude-User, PerplexityBot) in robots.txt, publish pages that answer one question outright in their opening lines, and build presence on the review sites, directories and comparison pages those assistants pull from. Rankings help with Google's AI Overviews; they matter far less to ChatGPT and Perplexity.
Why citation is a different game from ranking
A search result is a list. An AI answer is a recommendation. The list gives your prospect ten options and lets them choose; the recommendation gives them two or three and an implied endorsement. Getting named in the second is worth considerably more than placing eighth in the first.
The mechanics differ by engine, and the difference matters for where you spend your effort. Google is explicit about its own: to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet — the same foundations as classic SEO, with no additional technical requirement. ChatGPT and Perplexity work the other way round: they fetch and read pages themselves at the moment the question is asked, so a page that never cracked Google's first results can still be quoted verbatim to a buyer.
The practical consequence: if you only optimize for Google rankings, you compete for AI Overviews and ignore the assistants entirely. They are separate channels that happen to read the same web.
The discipline that covers the second channel has a name — Generative Engine Optimization, usually shortened to GEO. It is not a rebrand of SEO. It targets a reader that quotes rather than links, and it is judged on whether you get named, not on where you sit in a list. Google now runs two distinct AI surfaces of its own, AI Overviews and AI Mode, alongside the assistants — so "how to appear in AI search results" is really four questions, not one.
The crawler most businesses block by mistake
There are two families of AI crawler and they do completely different things. Training crawlers collect pages to build future models. Retrieval crawlers fetch pages live, at the moment a user asks a question, to ground the answer being written. Blocking one has nothing to do with blocking the other.
- Training —
GPTBot,ClaudeBot,Google-Extended. Blocking these keeps your content out of future model weights. It does not remove you from answers. - Retrieval and citation —
OAI-SearchBot,Claude-User,PerplexityBot. These are the ones that decide whether your name appears in an answer. Block them and you are invisible. - Neither, exactly —
Claude-SearchBotindexes for search quality rather than fetching at question time, andChatGPT-Userhandles user-triggered actions. OpenAI's documentation is blunt about the second one: "ChatGPT-User is not used to determine whether content may appear in Search."
The split is not symmetrical between providers, which is exactly why copied blocklists go wrong. OpenAI exposes two robots.txt levers and only OAI-SearchBot governs citation. Anthropic documents three bots, and the one that matters for being quoted is Claude-User — disabling it, in their words, "prevents our system from retrieving your content in response to a user query."
This is the single most common self-inflicted wound we see. A business decides it does not want its content used for AI training, copies a blocklist from a blog post, and takes out the retrieval bots along with the training ones. The intent was reasonable. The result is that no assistant can ever cite them.
One clarification worth having in writing, because it catches people out: Google-Extended does not remove you from AI Overviews. It governs training and grounding in Google's other systems. To limit what Search itself displays, Google points to the snippet controls — nosnippet, data-nosnippet and max-snippet — and those cut your ordinary search snippets at the same time.
Why third-party pages beat your own
Here is the finding that reorders most GEO plans. Seer Interactive analyzed 804,491 AI responses across 1,926 brands on four AI platforms. Review and trust sites came out as the second-largest citation source overall — and their share grows 10 to 20 times between the awareness stage and the moment someone is ready to buy. The closer your prospect gets to a decision, the more the answer leans on sources you do not own.
The same study graded brands into tiers built on Trustpilot profiles and found that the fully optimized ones picked up 9.5 times more co-mentions than brands with no verified active profile — meaning they get named when someone asks about a competitor. Read the multiplier for what it is: measured on one platform, and earned by a complete profile rather than by merely having one. It is still a competitive visibility channel you cannot buy on your own domain.
It reflects how these models weigh evidence: a claim about you on a page you control is an assertion, while the same claim on a platform where customers can contradict you is closer to a fact. The models behave like a cautious buyer, and a cautious buyer discounts the brochure.
If your entire visibility plan lives on your own domain, you are absent from the source category that carries the most weight exactly when the decision is being made.
None of this makes your own site pointless — it is where the detail lives, and it is what a curious buyer opens after the assistant names you. But it does change the order of work. Claim and complete the listings first, then make your own pages the place that resolves the details.
What actually decides who gets named
Strip away the tactics and three things determine whether a model can name you with confidence.
- Retrievability — the retrieval crawlers can reach the page, render it, and read the answer without executing your JavaScript to find it.
- Extractability — the answer to the question exists as a self-contained passage, not as a conclusion the reader is expected to assemble from six paragraphs.
- Corroboration — the same claim about you appears somewhere you do not control.
The second one is where most well-written content fails. Good marketing prose builds to its point. A retrieval system quotes a passage; it does not read your argument to the end. If someone asks "what does this company charge" and your pricing is implied across a page rather than stated in a sentence, there is nothing to lift. Answer the question first, then justify it.
The technical checklist
None of this is exotic. It is the same discipline as technical SEO, applied to a different reader.
- Audit
robots.txtfor retrieval bots specifically — confirmOAI-SearchBot,ChatGPT-User,Claude-SearchBotandPerplexityBotare allowed, whatever you decide about training. - Serve the substance server-side. If a passage only exists after hydration, assume it is not being read.
- Mark up what you are:
Organization,Service,FAQPagewhere genuine. Structured data does not force a citation, but it removes ambiguity about what your page is claiming. - Write one question per heading and answer it in the first two sentences underneath.
- Date your pages and keep the dates honest. A stale date on fresh content is worse than no date.
- Claim your listings on the platforms that serve your category, and get real reviews on them.
How to know whether it worked
You cannot check this in Search Console — AI assistant citations do not appear there. The workable method is unglamorous: write down the twenty questions a buyer would actually ask before choosing someone like you, ask each one across ChatGPT, Perplexity, Gemini and Claude, and record who gets named. That is your baseline. Repeat it monthly.
Two warnings, both learned the hard way. Answers vary between runs for the same prompt, so a single check proves very little — run each question more than once before you conclude anything. And personalized or logged-in sessions skew results; use a clean session.
Frequently asked questions
Does blocking GPTBot stop ChatGPT from recommending my business?
It stops OpenAI from using your pages as training data, but it does not stop ChatGPT from citing you. Citation is governed by a different crawler: OAI-SearchBot. OpenAI states that ChatGPT-User is not used to determine whether content may appear in Search, so OAI-SearchBot is the one to allow. On the Anthropic side the equivalent is Claude-User, which retrieves your content when someone asks Claude a question. Many sites block the training bot and assume they are safe, without realizing the retrieval bots are a separate decision.
Do I need to rank on Google page one to be cited by AI?
It depends on the engine. For Google AI Overviews and AI Mode, Google states that a page must be indexed and eligible to be shown in Search with a snippet — the same foundations as classic SEO, with no additional technical requirement. ChatGPT and Perplexity retrieve pages themselves at question time, so a page that never ranked on Google can still be quoted. The two channels reward different things and need separate work.
Is there a way to opt out of AI Overviews specifically?
Not on its own. Google-Extended controls training and Gemini grounding, but it does not remove you from AI Overviews. What does apply is snippet control: nosnippet, data-nosnippet and max-snippet limit what Google can display, including in AI Overviews. The trade-off is that the same directives also shrink your regular search snippets.
How long before AI assistants start naming us?
Slower than a ranking change and faster than a domain aging. Retrieval-based engines re-crawl and re-index continuously, so a new page can be quoted within weeks. Building the third-party presence that makes you a default recommendation — directories, review platforms, mentions on sites the models already trust — takes months. Anyone promising a specific date is guessing.
If you want this done rather than explained, that is our job — we run the baseline, fix the technical blockers, and build the off-site presence that makes the citations stick. Book a call and we will tell you where you currently stand.
