Home / Blog / AI search audit
How-to

The AI search visibility audit: a self-serve checklist before you buy a tool

Before you pay for an AI visibility platform, spend an afternoon auditing yourself by hand. It teaches you the mechanics no dashboard will: how a crawler reaches your pages, why an answer engine names one brand over another, and where you quietly vanish from a buyer's shortlist. This scored, five-part checklist shows you exactly where you stand across ChatGPT, Gemini, and Perplexity, and why brands automate the moment the first pass is done.

[ AI VISIBILITY AUDIT ]Audit your AIanswer visibility71,000:1crawl requestsper referral, AnthropicCloudflare Radar, week of 19 June 2025A scored, do-it-yourself checklist you can run in an afternoon.
Score five sections out of 100 to find where you disappear from AI answers.

Why a manual audit beats a dashboard on day one

Buying an AI visibility tool before you understand your own exposure is like hiring a personal trainer before you have stepped on a scale. You will get numbers, but you will not know which ones matter or why they moved. A first-pass audit you run by hand teaches you the mechanics: how a crawler reaches your pages, how an answer engine decides to name you, and where your brand quietly disappears from a buyer's shortlist. That understanding is what makes a tool worth paying for later, because you will know exactly what you are automating and what "good" looks like.

The audit below is scored out of 100. It has five sections worth twenty points each: crawler access, multi-platform prompt testing, citation-pattern analysis, content structure, and schema. Work through it once for your own domain and two or three competitors. It takes an afternoon. The point is not the number itself but the gaps it exposes, and the reason those gaps are almost impossible to keep closed by hand.

Section one: crawler access (20 points)

If the models cannot fetch your pages, nothing else in this audit matters. Open your robots.txt (it lives at yourdomain.com/robots.txt) and check it line by line against the current AI user-agents. There are more of them than most teams realise, and they split into two jobs. Training crawlers gather content to improve future models. Retrieval crawlers fetch a page in real time to answer a live question, and those are the ones that produce citations.

Score the full twenty only if every retrieval agent is explicitly allowed and you have made a deliberate, documented choice on each training agent. A common own goal: a blanket "Disallow: /" left over from a staging config, or a security plugin that silently blocks unknown user-agents. Note that blocking a training crawler is a legitimate business decision and does not, on its own, remove you from AI answers, because retrieval and training are separate pipelines. But blocking a retrieval agent will quietly erase you from that platform's citations.

Curious how AI engines describe your brand right now? Get a free visibility audit and see where you stand across ChatGPT, Gemini and Perplexity.

Section two: multi-platform prompt testing (20 points)

This is the heart of the audit and the part almost everyone skips. Write down the ten to fifteen questions a real buyer would ask on their way to a purchase. Not "what is your brand", but the category questions: "best AEO platform for a B2B SaaS", "how do I get cited by Perplexity", "alternatives to X for enterprise". Then ask each question, verbatim, in ChatGPT, Gemini, and Perplexity. Google AI Overviews and Copilot are worth adding if they matter to your market.

For every prompt on every platform, record three things: were you named, were you cited with a link, and who was named instead. Use a simple grid. A cell gets full marks if you are named and linked, half marks if you are mentioned without a link, and zero if you are absent. This is tedious and it is meant to be, because the tedium is the lesson. The output is a heatmap of where you win and where a competitor owns the answer.

"If you are not in the answer, you are not in the consideration set - and the buyer never learns you existed."

Two mechanics make this harder than it looks. First, answers are non-deterministic: ask the same question twice and the wording, and sometimes the brands named, will shift. A single test is an anecdote, not a measurement, so run each prompt a few times. Second, personalisation and memory skew results, so test in a logged-out or fresh session. Score this section on coverage: full marks if you are named in the majority of your buyer prompts across at least two platforms.

Section three: citation-pattern analysis (20 points)

Being cited is good. Understanding why you were cited is what lets you do it on purpose. For every answer where a competitor was named instead of you, click through to the source the engine used. You are looking for the shape of the winning page. Was it a listicle, a comparison table, a documentation page, a Reddit thread, a review site? Answer engines lean heavily on third-party and structured sources, not just brand-owned pages.

Here the overlap with classic SEO is real but partial. A Semrush analysis of 5,000 queries found Google's AI Mode shared roughly 51% domain overlap and about 32% URL overlap with the top ten organic results, while Google AI Overviews aligned far more closely at around 86% domain and 67% URL overlap. The practical reading: strong organic rankings help, especially for AI Overviews, but a meaningful share of AI citations comes from pages that do not rank in the classic top ten. Ranking is necessary insurance, not a guarantee.

Log the pattern for each lost answer. If three of your buyer questions are won by a third-party "best tools" listicle you are not on, that is a concrete, fixable finding: get accurately represented on that page. If the winning source is a competitor's own comparison page, that tells you which asset you are missing. Score this section on how completely you can explain your losses. Full marks if, for every gap, you can name the source type and the specific page that beat you.

Section four: content structure (20 points)

Models extract answers, they do not read prose the way a human does. Pages that get quoted tend to answer a question directly, near the top, in a self-contained passage a model can lift without stitching together. Walk your key pages and score them for extractability.

The tell-tale failure is a page that ranks fine for humans but gives a model nothing clean to quote: a great answer buried in paragraph nine, wrapped in narrative. Score twenty if your priority pages consistently lead with the answer and break cleanly into quotable chunks.

Section five: schema and machine-readability (20 points)

Structured data does not force a citation, but it removes ambiguity about what your page is and who you are. Check that your key pages carry valid schema.org markup: Organization and, where relevant, Product, FAQPage, Article, and Breadcrumb. Validate it, because broken JSON-LD is worse than none. Confirm your Organization markup states your name, logo, and sameAs links to your authoritative profiles, so an engine can resolve "who is this" without guessing.

Machine-readability goes beyond schema. Ensure primary content is present in the server-rendered HTML rather than painted in by client-side JavaScript that a lightweight retrieval fetch may never execute. A retrieval crawler that gets an empty shell will cite whoever gave it clean text instead. Score this section on validity and coverage across your priority templates, not on having thrown schema at the homepage alone.

Add up the score, then read the real lesson

Total the five sections. Under 50 and you have structural problems keeping you out of answers entirely. Between 50 and 75, you are visible in patches and losing winnable questions to better-prepared competitors. Above 75, you are in good shape and the game becomes defending and widening your lead. Whatever the number, you now have a ranked list of specific, concrete fixes rather than a vague sense that "we should do something about AI".

Here is the honest part. Everything above is a snapshot, and the snapshot decays within days. Answers are non-deterministic, so your prompt grid shifts every time you rerun it. Models update, competitors publish, and a page that named you last Tuesday may not this Friday. The economics of the underlying crawling make this vivid: over one week in June 2025, Cloudflare measured Anthropic's crawler making close to 71,000 page requests for every visitor it referred back, a reminder that these systems ingest the web at a scale and speed no manual pass can track. The web an engine answers from today is not the one it answered from last month.

That is the natural ceiling of a hand-run audit, and also its value. Do it once, properly, and you learn the mechanics well enough to know what to watch. But watching fifteen prompts across five platforms, several times each, for yourself and your competitors, week after week, is a monitoring problem, not a project. It is the point at which brands stop doing this in a spreadsheet and start automating it: continuous prompt testing, citation tracking over time, and alerts when you drop out of an answer you used to own. The manual pass tells you where you stand. A system tells you the moment that changes.

See where you stand across every answer engine

Run the manual audit once, then let Stellarcast keep watch. We track whether ChatGPT, Gemini, Perplexity, Google AI Overviews, and Copilot name and cite your brand across your real buyer questions, and alert you the moment you drop out of an answer you used to own.

Get your free visibility audit

Frequently asked questions

Do I need to allow AI crawlers in robots.txt to appear in ChatGPT or Perplexity?

For live citations, yes. Retrieval agents such as ChatGPT-User, Claude-User, and Perplexity-User fetch your page in real time to answer a question, and blocking them removes you from that platform's cited answers. Training crawlers like GPTBot, ClaudeBot, and Google-Extended are separate: blocking them stops your content feeding future models but does not, by itself, remove you from answers. Check each token individually rather than using a blanket rule.

How many buyer questions should I test in the audit?

Aim for ten to fifteen category questions that map to a real buying journey, not brand-name lookups. Test each one across at least ChatGPT, Gemini, and Perplexity, and run every prompt a few times because answers are non-deterministic. That gives you enough coverage to see genuine patterns rather than a single lucky or unlucky response. Fewer than ten prompts and you risk drawing conclusions from noise; the grid is only useful once it is broad enough to reveal where you consistently win or lose.

Does ranking well on Google mean I will get cited in AI answers?

It helps but does not guarantee it. A Semrush study found Google AI Overviews overlapped heavily with organic results, around 86% at the domain level, while the newer AI Mode shared only about 51% domain and 32% URL overlap with the top ten. So a meaningful share of AI citations comes from pages outside the classic top ten, including third-party listicles, review sites, and forums. Strong rankings are useful insurance, not a complete strategy for AI visibility.