How Stellarcast measures AI visibility
We write a lot about how AI engines decide who they name. It is only fair to be just as open about how we measure it. This is our methodology in plain terms: which engines we query, how often, what we read from each answer, and the rules we hold ourselves to, including the ones that stop us from handing you a confident number that is quietly wrong. If you are going to trust our numbers, you should know exactly how they are made.
The loop: monitor, diagnose, execute, prove
Measurement is not the whole product, but it is the foundation the rest stands on. Our work runs as a loop: monitor how visible a brand is across AI engines, diagnose the specific gaps that hold it back, execute the fixes, and then prove whether visibility actually moved. Everything in this article is about the first and last steps, monitoring and proving, because those are the parts that live or die on honest measurement. If the monitoring is soft, the diagnosis is guesswork and the proof is theatre.
That framing shapes every decision below. We would rather report a wider, truer band than a narrow, flattering point. We would rather tell you an engine answers from training data with no live web than lump it in with the ones that cite real pages. The rules that follow are what keep the loop honest.
The engines: a core set and an extended set
We measure nine engines, and we do not pretend they all deserve the same attention or work the same way. We split them into two tiers by how much they matter to real buyers and how much each measurement costs to run.
- Core engines, measured on the fastest cadence: Google AI Overviews, ChatGPT, Gemini and Perplexity. These carry the bulk of real AI-answer traffic, so they are the ones we re-read most often.
- Extended engines, measured less frequently: Google AI Mode, Claude, Copilot, Grok and DeepSeek. These matter, but they are lower-traffic, so we re-read them on a slower cadence rather than daily.
The split is not cosmetic. Reading the highest-traffic engines most often, and the rest on a slower cadence, is what lets us cover nine engines at all rather than four. When you read a Stellarcast number, it comes with the engine attached, because a blended score across engines you do not care about is one of the easiest ways to be misled.
Honesty about what each engine actually is
Not every "AI engine" answers the same way, and flattening that difference is a quiet form of lying. Some engines answer from the live web and return citations you can inspect. Others answer from training knowledge alone, with no live retrieval and no sources. DeepSeek, for example, answers from training knowledge with no live web and no citations, so we label it exactly that way rather than implying it browsed the internet to reach its answer. Google AI Overviews, ChatGPT search, Gemini and Perplexity behave more like live, citing systems; Copilot draws on a search index behind it. We keep those distinctions visible in the data, because "you were not mentioned" means something very different on a live-web engine than on one answering from a year-old training snapshot.
"A number without a spread hides the noise, it does not remove it. So we report a band, not a single figure."
Want to see this run against your brand? Get a free visibility audit and we will show you the real per-engine picture.
The prompts: a fixed, native, per-locale panel
Everything starts with the questions. We build a query universe from the real, buyer-style questions people ask in a category, expanded across the topics a brand wants to watch, and we keep the panel fixed so that runs stay comparable over time. Change the questions and you change the measurement, so the panel is treated as a controlled instrument, not a thing to tweak between runs.
Crucially, prompts are native to each locale, not machine-translated from one base market. A German buyer's question is written in German the way a German buyer would phrase it, because a translated prompt measures a translated question, not the real one. Locale is a first-class part of every measurement we take, not an afterthought bolted on at the end.
What we read from each answer
Once an answer comes back, we extract a small number of things that actually matter, and we compute them per engine and per locale rather than as one global blur.
- Visibility rate. Whether your brand is mentioned in the answer at all, expressed as how often you appear across the panel.
- Mention-share of voice. Your mentions as a share of all brand mentions in the same answers, measured against your confirmed rivals. This is the "are you in the conversation, and how big is your slice" number.
- Citation share of voice. Which pages the engine actually cited, resolved to a URL and domain. Presence is one thing; being the source the answer is built from is another.
- Owned, earned or rival. Every citation is classified as one of yours, an earned third-party page that mentions you, or a competitor's. That tells you whether the answer is standing on your own material, on independent corroboration, or on a rival's page.
Rivals are not guessed by string-matching a name. Competitors are confirmed entities, so a "rival cited while you are absent" signal means a real, named competitor, not a coincidence of wording.
Why we report a band, never a bare number
This is the rule we are most stubborn about. AI answers are non-deterministic: ask the same engine the same question twice and you can get two different answers. A single "your visibility is 34 percent" would imply a precision that does not exist. So we do not show a bare number. We show a range across repeated runs, so you can see the spread, a percentile band rather than a single point. A tight band means a stable result you can act on; a wide one means the engine is genuinely inconsistent about you, which is itself something worth knowing. Hiding that behind one clean figure would be the most flattering thing we could do and the least honest.
The cadence: measurement is a signal, not a project
Citation patterns move, sometimes sharply, within weeks. A source that dominates an engine's answers this month can be a minor player the next. That is why we re-measure on a schedule rather than running a one-off audit and calling it done. Core engines are re-read on the fastest cadence, extended engines less often, and every reading is stored as an append-only, date-stamped time series with the engine's model version and the prompt version attached to each row. Nothing is silently overwritten, so when a number moves, we can trace exactly what changed and when. A measurement without a date, in this space, is barely a measurement at all.
Proving lift honestly, including when we cannot
The hardest discipline is the last one: saying whether a change actually worked. When a brand makes a fix, we compare visibility before and after, but we refuse to overclaim. Every lift we report carries a confidence label. A result is marked measured only when the evidence is strong enough, likely when it points that way but is not yet conclusive, and too-early when not enough time or data has passed to say anything at all. "Too early" is a real, common answer in our system, not a failure. Telling you a change worked before the data can support it would be exactly the kind of confident-but-wrong claim this whole methodology exists to prevent.
The rules, in one place
If you take nothing else from this, take the short list of commitments the measurement holds itself to. They are the same tests we suggested you apply to any tool, including ours.
- Ranges, not single numbers. Every visibility figure is a band, so the noise is visible.
- Per engine, always. Nine engines, read separately, never blended into one flattering average.
- Per locale, always. A native prompt panel per market, because a translated question is a different question.
- Sources named and classified. Every citation resolved and marked owned, earned or rival.
- Everything date-stamped and append-only. Because a citation snapshot without a date is already going stale.
- Confidence on every claim of lift. Measured, likely or too-early, and we say too-early when it is true.
The takeaway
Good measurement in AI search is mostly a set of refusals: refusing to show a single number when the truth is a range, refusing to blend engines that behave nothing alike, refusing to translate a prompt and call it the local question, and refusing to claim a lift before the data earns it. That is the discipline behind every figure we publish, and it is why we are comfortable being this open about the method. If a number is worth showing you, it is worth showing you how it was made.
See your real per-engine picture
This is the method behind every Stellarcast audit: nine engines, every locale, ranges not guesses, and citations classified as yours, earned or a rival's. Request a free visibility audit and see where you actually stand.
Get your free visibility auditFrequently asked questions
Which AI engines does Stellarcast measure?
Nine. We split them into core engines, measured on the fastest cadence because they carry the most traffic, and extended engines, measured less often. The core set is Google AI Overviews, ChatGPT, Gemini and Perplexity. The extended set adds Google AI Mode, Claude, Copilot, Grok and DeepSeek. We label each engine honestly: some answer from the live web with citations, while others, such as DeepSeek, answer from training knowledge with no live web, and we mark them as such rather than pretending every engine works the same way.
Why does Stellarcast show a range instead of a single visibility number?
Because AI answers are non-deterministic: the same question does not reliably return the same answer. A single clean number would be false precision. Instead we report a band, showing the spread across repeated runs rather than one figure, so you can see how stable or noisy a result really is. A number without a spread hides the noise, it does not remove it.
How does Stellarcast measure share of voice?
We run a fixed set of real, buyer-style prompts against each engine on a schedule, then read the answers for two things: whether your brand is mentioned, which gives a visibility rate and a mention-share of voice against your rivals, and which pages are cited, which gives a citation share of voice. Every result is computed per locale and per engine, date-stamped, and rolled up over a time window so the trend, not any single run, is what you read.
How are cited sources classified?
Every citation an engine makes is resolved to a URL and domain and classified as owned, meaning your own properties, earned, meaning third-party pages that mention you, or a rival's. That is what lets us tell you not just whether you appear, but whether the answer is being built from your own material, from independent corroboration, or from a competitor's page instead.
How often does Stellarcast re-measure?
Continuously, on a schedule, because citation behaviour shifts within weeks. Core engines are re-measured on the fastest cadence and extended engines less frequently. Every reading is stored as an append-only, date-stamped time series with the engine model and prompt version attached, so a change in the numbers can always be traced and nothing is silently overwritten.