How to measure your brand's visibility in AI answers
Before you can improve how AI describes your brand, you have to see it clearly. Measurement is the first move - and it's more than typing your name into ChatGPT once and nodding.
Most teams discover their AI visibility by accident: someone asks an assistant for a recommendation in their own category, watches a competitor get named, and feels the floor shift. That moment is useful, but it's an anecdote. To act on it you need a method - one that turns "I think we're invisible on Perplexity" into "we're named in 2 of 10 buying questions on Perplexity, versus 7 of 10 for our top competitor, and here's where."
Why measurement comes first
AI answers are invisible by default. There's no dashboard from the engines telling you how often you're named, no rank tracker built for citations. If you skip measurement and jump straight to "publishing more content," you're guessing - and you'll have no way to know whether anything you did moved the needle. A baseline is what converts effort into evidence.
Curious how AI engines describe your brand right now? Get a free visibility audit and see where you stand across ChatGPT, Gemini and Perplexity.
What to actually measure
Five signals capture most of what matters. Track each one per engine, because the same question yields different answers on different assistants.
- Mention rate - across your set of buying questions, how often is your brand named at all?
- Citation rate - when you're mentioned, is a source linked back to you (your site, a review, an article)? Citations are stickier than passing mentions.
- Share of voice - of the brands named for a question, what share are you versus each competitor? This is the number that reframes the conversation internally.
- Factual accuracy - when the model describes you, is the category, pricing, and feature set correct? A confident-but-wrong description is its own problem.
- Framing and sentiment - are you positioned as the leader, a budget option, a niche pick? The adjectives matter as much as the mention.
Run a manual baseline first
You can start by hand in an afternoon. It won't scale, but it teaches you what to look for.
- Write your buying questions. Not "what is [your brand]" - the questions a buyer asks when they're choosing: "best [category] for [use case]", "[competitor] alternatives", "is [your brand] good for [segment]". Aim for 15-25 that reflect real intent.
- Ask every engine. Run each question on ChatGPT, Claude, Perplexity, Gemini and Copilot. Use a fresh session so prior chat doesn't skew the answer.
- Record the answer, not your impression. For each question and engine, note: were you named? cited? who else appeared? was the description accurate? Put it in a simple grid.
- Tally the patterns. Now you can see it - the questions you own, the ones a competitor owns, and the engines where you're absent entirely.
Skip the spreadsheet
Stellarcast runs this continuously: it asks the engines your buyers' real questions, tracks mention rate, citation rate and share of voice against competitors, flags when the facts about you are wrong, and re-checks after every change so you can prove the lift. Request a free audit to see your baseline.
Get your free visibility auditWhy a one-off check isn't enough
A manual audit is a photograph; AI visibility is a film. Engines update their models and re-crawl their sources on their own cadence, competitors publish, and a fact that was right last month goes stale. An answer you screenshotted in March may be different in June - for better or worse - and you'd never know. The value of measurement compounds only when it's repeated on the same questions over time.
Turn measurement into a baseline you can act on
The point of measuring isn't a number on a slide - it's a map of where to work. A good baseline tells you three things: which buying questions you're losing, which engines you're weakest on, and which competitors are eating your share. From there the path is concrete:
- Diagnose the missing facts, sources or entities behind each absence.
- Fix the source of truth so models can cite you confidently.
- Re-measure the same questions and tie each change to a real movement in how AI names you.
That last loop - measure, change, measure again - is the whole game. Without it you're publishing into the dark. With it, every improvement is something you can prove.
"A prompt set that only asks "what is [your brand]" measures your fan club, not your market."
Build a prompt set that mirrors the buying journey
Your prompt set is the instrument. If it's biased, every number downstream is biased too. The most common mistake is stuffing it with brand-aware questions - "what is [your brand]", "[your brand] pricing" - which the model answers well because it's reading your own site. That measures your fan club, not your market. A useful set is weighted toward the questions buyers ask before they know your name.
Structure it by intent stage so your coverage isn't lopsided:
- Awareness (problem-first): "how do I stop [pain point]", "why is [process] so slow". No brand or category assumed. This is where you're most likely to be absent and most valuable to appear.
- Consideration (category-first): "best [category] for [segment]", "top [category] tools 2026", "[category] with [must-have feature]". The battleground for share of voice.
- Decision (comparison-first): "[competitor] alternatives", "[your brand] vs [competitor]", "is [your brand] worth it for [use case]". Where accuracy and sentiment decide the deal.
Aim for roughly a third in each stage. Pull the exact wording from real sources rather than inventing it: your sales team's call notes, the "people also ask" boxes on your category terms, and your own site search logs. Phrasing matters more than you'd think - "affordable [category]" and "enterprise [category]" surface completely different brand sets from the same engine.
Keep the set versioned. When you add or reword a prompt, note the date, because your trend line only means something if the questions stayed constant between readings. Treat a prompt-set change like a schema migration: deliberate, logged, and never silent.
Turn the five signals into numbers you can defend
"We show up sometimes" doesn't survive a leadership meeting. Convert each signal into a rate so it's comparable across engines and across months. The math is deliberately simple - the discipline is in counting consistently.
- Presence rate = questions where you're named / total questions asked. Run per engine. If you're in 4 of 20 consideration prompts on Perplexity, that's 20 percent, and you can watch it move.
- Share of voice = your mentions / all brand mentions across the set. If a typical answer names five brands and you're one of them, your ceiling is telling: even perfect presence caps you near 20 percent unless answers get shorter or you displace a competitor.
- Citation share = answers that link to your domain / answers that link to anyone. This is distinct from mentions. A brand can be named often yet cited never, which means the model likes your reputation but pulls its facts from someone else's page. Track the two separately or you'll misdiagnose the fix.
- Accuracy rate = correct descriptions / total descriptions of you. Log the specific errors (wrong category, stale pricing, a feature you sunset) because those become your content backlog.
- Sentiment = a coded label per answer (leader / solid / niche / budget / negative). Resist a numeric average here; the distribution is what you act on. Ten "budget option" tags is a positioning problem, not a rounding error.
Record raw counts alongside the rates. When someone challenges the number in a review, you want to reopen the underlying grid and point at the exact answers, not defend a percentage in the abstract.
Set a cadence and sample enough to trust the number
A single run of your prompt set is a snapshot with real noise in it. Assistants sample their responses, so the same question can name different brands on two consecutive asks - response variability rises with the model's sampling temperature, which is why one run is an anecdote and three runs is a signal (as reported in work on LLM sampling temperature and output diversity). Ask each prompt on each engine at least three times and record the modal answer, or the fraction of runs you appeared in. That fraction is a more honest presence number than a lucky single hit.
Then pick a rhythm and hold it:
- Monthly for the full set on all engines. This is your trend line and your board number. Same prompts, same day of the month, same fresh-session hygiene.
- Weekly for a small "canary" subset - your ten highest-value consideration and comparison prompts. Enough to catch a sudden drop or a factual error going viral before the monthly run would.
- Ad hoc after any known trigger: a launch, a pricing change, a competitor's funding announcement, or a model version bump on one of the engines.
Run every engine every time, and never average them into one score. Engines disagree hard on sourcing - a 2026 audit reportedly found only about 11 percent overlap between the domains ChatGPT cites and those Perplexity cites - so a blended number hides exactly the per-engine gap you're trying to fix. Keep the columns split.
Turn each finding into a triaged next move
Measurement earns its keep only when a row in your grid becomes a task. Once the baseline is in, sort your findings into four buckets and act on them in order of leverage, not alphabetically:
- Wrong beats absent. A confident, inaccurate description is your top priority - it actively costs you deals. Trace which source the engine leaned on (the cited domains tell you), then correct the record at the source: your own page, a review site, a Wikipedia-adjacent reference the models trust. Re-measure that prompt next week to confirm the fix propagated.
- Absent where you should win. Consideration prompts where competitors appear and you don't. These need genuinely useful content that answers the exact question - a comparison page, a use-case guide - structured so an engine can lift and cite it. This is your slowest lever; start it now.
- Named but not cited. You're in the answer, but the link points elsewhere. Here the gap is a citable asset, not awareness. Publish the specific data, definition, or comparison the engine is currently borrowing from someone else and attributing to them.
- Sentiment mismatch. You appear and you're accurate, but framed as the budget or niche pick when you're neither. This is a messaging and third-party-proof problem - case studies, analyst mentions, the language on the pages the model reads.
Assign each bucket an owner and a re-measurement date, and log the "before" number next to it. The next monthly run isn't just a fresh snapshot then - it's a scorecard that tells you which moves worked and which didn't touch the number at all.
Frequently asked questions
How do I check if my brand appears in ChatGPT or Perplexity?
Ask each engine the real questions your buyers ask - the ones where someone is choosing a product or vendor - and record whether you're named, whether a source is cited back to you, and which competitors appear instead. Repeat across engines, since results differ.
What metrics matter for AI visibility?
Mention rate (how often you're named), citation rate (how often a source links back to you), share of voice versus competitors, factual accuracy of what's stated about you, and framing or sentiment. Track each per engine.
Why do AI engines give different answers to the same question?
Each uses different models, training data, retrieval sources and refresh cadences. The same prompt can name different brands on ChatGPT, Perplexity and Gemini - which is why visibility has to be measured per engine, not assumed.
How often should I measure?
A one-off audit shows where you stand today, but answers shift as sources and models update. Continuous tracking - re-running the same prompts on a schedule - is what lets you tie a change you made to a real movement in how AI describes you.