Short answer
Run the audit in five steps, in this order: check that crawlers can fetch and render your pages; check that each page answers one real buyer question; measure how often you are named across the assistants your buyers use; check how you are described on sites you do not own; then rank the findings by what has to be true before anything else can work. Two findings override all others — a retrieval crawler being blocked, and content that only exists after JavaScript runs. Everything else is optimization on top of a site that already works.
What an audit is supposed to decide
An audit that produces a report and no decision was theater. Before running anything, write down the decision it has to serve. In almost every case it is one of three: what do we fix first, what do we stop doing, and what do we need to be able to see next quarter.
That framing kills most of what fills a standard audit deck. A list of two hundred pages with a missing meta description is not a decision. It is a spreadsheet that survives one meeting and then gets ignored, because nobody can tell which of the two hundred lines mattered. If the distinction between ranking in a list and being cited inside an answer is new to you, start with what generative engine optimization actually is — the method below assumes it.
A finding earns a place in the report only if you can complete this sentence about it: "if we do not fix this, then X cannot happen." If the sentence has no end, the finding is trivia. This single filter usually cuts an audit down by three quarters, and makes the remainder actionable.
Step 1 — can machines read the site at all
Everything downstream assumes this. Start here, and stop the audit if it fails, because there is no point auditing content that nothing can fetch.
- Fetch your
robots.txtand read it line by line. Look specifically for retrieval crawlers:OAI-SearchBot,Claude-User,PerplexityBot. These govern whether you can be quoted. Training crawlers such asGPTBotandClaudeBotare a separate policy question with no effect on citation. - View the rendered page with JavaScript disabled. If the article body disappears, you have found the most expensive problem on the site. Content that only exists after hydration may never be read.
- Check status codes on the pages that matter, not on a sample. A commercial page returning a soft 404 or an unintended redirect chain is common after a migration and invisible from the browser.
- Confirm the sitemap matches reality. Every live page in it, no dead URLs, no pages you deliberately keep out of the index.
- Check for
nosnippetandmax-snippetdirectives. They apply to Google's AI features too, so a snippet restriction set years ago for a different reason now removes you from AI Overviews — the eligibility rules, and the control that does not do what its name suggests, are in how to appear in AI Overviews and AI Mode.
The crawler distinction in the first bullet is the one that gets written up wrong most often, including by agencies. We set out which bot controls what, with the vendor documentation behind each claim, in how to get cited by ChatGPT and Perplexity.
Step 2 — does the content answer real questions
This step is a comparison, not an inspection. On one side, list the questions your buyers actually ask — from sales calls, support tickets, and the last ten minutes before a contract is signed. On the other, list the pages you have published. Then line them up.
Three patterns show up every time, and each has a different fix.
- Questions with no page. The gap that costs the most, and the easiest to act on. Usually these are the late-stage questions — pricing logic, integration constraints, what happens when it goes wrong.
- Pages with no question. Written because a keyword tool suggested them. They dilute the site and are candidates for retirement or consolidation.
- Several pages on one question. Splits the signal across URLs and leaves a retrieval engine choosing between two of your own pages. Merge into the strongest URL and redirect the rest.
For the pages that survive, apply one structural check: is there a self-contained answer near the top that could be lifted out and quoted alone? A page whose substance arrives in paragraph five offers a retrieval system nothing to extract. The content strategy that serves both SEO and AI search covers how to bake that into the brief rather than retrofit it page by page.
Step 3 — the citation baseline
This is the step nobody wants to do manually and the one that produces the only number worth reporting to a board.
Run the twenty-question method we set out in how to get cited by ChatGPT and Perplexity, using questions from step 2 rather than generic ones. One addition matters when you are auditing rather than tracking: record a third column next to "were we named" and "who else was named" — what source the answer leaned on.
That third column is what turns a measurement into a finding, and it is the column trackers do not give you. If the same third-party site keeps appearing as the source across your category, you have just located where the next quarter of off-site work belongs, and step 4 has half its answer before you start it.
Audit-specific caveat: do not build a baseline the week of a migration or a redesign. You will attribute months of movement to a change that was actually the site being briefly unfetchable, and the baseline becomes unusable as a comparison point for everything that follows.
Search Console still matters alongside this. Google's documentation states that sites appearing in AI features are included in overall search traffic, reported under the "Web" search type rather than broken out separately — the implications for reading a traffic drop are covered in how to measure a plan that has two outcomes. Treat that report as your search picture, and the manual baseline as your AI picture.
Step 4 — how you are described off your own site
A claim on your own site is an assertion. The same claim on a site you do not control is closer to evidence, and models weight it accordingly. So the audit has to leave your domain.
- Search your company name and read what comes back — directories, review platforms, old profiles, an outdated address on a listing nobody has logged into for three years.
- Check consistency of the basics across every profile you find: what you do, who you serve, where you are. Contradictions between sources are exactly the uncertainty that makes a model reach for a better-documented competitor instead.
- Note where you are absent from the review platforms and directories your competitors occupy, especially the ones that surfaced as sources in step 3.
- Look at what people say, not only whether they say it. Review platforms are among the sources AI answers lean on most in commercial categories, and the way you collect reviews is legally constrained in ways worth checking before you scale it.
Step 5 — turning findings into an ordered list
Findings without an order are a to-do list that nobody starts. Order them by dependency first and effort second — what has to be true before anything else can work.
- Blockers. Retrieval crawlers denied, body content that requires JavaScript, broken status codes on commercial pages. Nothing else you do has any effect until these are cleared.
- Structural fixes. Merging duplicate pages, retiring dead ones, adding an extractable answer to the pages that already attract traffic. Cheap, and they compound.
- Gaps. New pages for questions with no page. Slower to pay off, and this is where the content calendar comes from.
- Off-site. Profiles, listings, review presence. Slowest of all because it depends on other people publishing, which is precisely why it should start early rather than last.
Put a named owner and a date on each item before the report is circulated. An audit whose findings have no owner produces exactly as much change as no audit at all, and costs more.
Three traps in bought audits
If you are commissioning one rather than running it, three things separate a diagnosis from a deliverable.
- A score out of 100. Composite scores average unrelated things — a blocked retrieval crawler and a missing alt attribute land in the same number. Ask instead which findings are blockers and which are not.
- An automated AI visibility report with no methodology. Ask which questions were asked, how many times each, and on which assistants. If the answer is vague, the number is not reproducible and cannot be compared next quarter.
- Recommendations that cannot be refused. "Publish more content" and "improve E-E-A-T" survive any evidence. A real finding names a page, a cause and a fix.
Frequently asked questions
How long does an AI SEO audit take?
A focused one takes two to three days of work for a site under a few hundred pages, and most of that is the citation baseline — asking twenty questions across four assistants and recording the answers is slow manual work. The technical checks take an afternoon. Anything sold as a fifteen-minute automated audit is producing a score, not a diagnosis, because the part that matters cannot be crawled.
Which tools do I need to run one?
Google Search Console, a crawler that can render JavaScript, your browser, and a spreadsheet. That is genuinely it for a first pass. Paid AI visibility trackers automate the citation baseline, which saves time once you already know which questions matter — but they cannot tell you which questions matter, and that is the decision the audit exists to make.
Should I block AI crawlers while I fix things?
Almost never, and the distinction matters because the two kinds of bot are separate. Blocking a training crawler such as GPTBot or ClaudeBot has no effect on whether you are cited. Blocking a retrieval crawler such as OAI-SearchBot, Claude-User or PerplexityBot removes you from the answers entirely, and recovery is not instant once you unblock. Fix the pages while they stay accessible.
How often should the audit be repeated?
The technical section once or twice a year, or after any migration, redesign or framework change — those are what silently break rendering and status codes. The citation baseline monthly, because it is the only number that tells you whether the work moved anything, and because a single measurement of it is unreliable on its own.
If you would rather have the five steps run for you — including the manual baseline, which is the part that takes the time — book a call and we will tell you which of the two blockers you have before we quote anything.
