Home / Blog / How Stellarcast measures AI visibility
Methodology

How Stellarcast measures AI visibility

We write a lot about how AI engines decide who they name. It is only fair to be just as open about how we measure it. This is our methodology in plain terms: which engines we query, how often, what we read from each answer, and the rules we hold ourselves to, including the ones that stop us from handing you a confident number that is quietly wrong. If you are going to trust our numbers, you should know exactly how they are made.

[ OUR MEASUREMENT LOOP ]Nine engines. Every locale.Ranges, not single numbers.MONITORDIAGNOSEEXECUTEPROVECore engines daily. Extended engines on a slower cadence.Every reading date-stamped, per locale, stored append-only.Cast your brand across AI - and measure whether it landed.
Our loop: monitor visibility across engines, diagnose the gaps, execute fixes, and prove the lift. Every reading is per locale, date-stamped, and stored as an append-only time series.

The loop: monitor, diagnose, execute, prove

Measurement is not the whole product, but it is the foundation the rest stands on. Our work runs as a loop: monitor how visible a brand is across AI engines, diagnose the specific gaps that hold it back, execute the fixes, and then prove whether visibility actually moved. Everything in this article is about the first and last steps, monitoring and proving, because those are the parts that live or die on honest measurement. If the monitoring is soft, the diagnosis is guesswork and the proof is theatre.

[ IT NEVER STOPS — IT LOOPS ] MONITOR visibility DIAGNOSE the gaps EXECUTE the fixes PROVE the lift per engine per locale Every reading date-stamped and stored append-only, so the loop can prove what changed.
The loop runs continuously: monitor visibility across engines, diagnose the gaps, execute the fixes, prove the lift, then measure again. Every reading is per engine, per locale, date-stamped and append-only.

That framing shapes every decision below. We would rather report a wider, truer band than a narrow, flattering point. We would rather tell you an engine answers from training data with no live web than lump it in with the ones that cite real pages. The rules that follow are what keep the loop honest.

The engines: a core set and an extended set

We measure nine engines, and we do not pretend they all deserve the same attention or work the same way. We split them into two tiers by how much they matter to real buyers and how much each measurement costs to run.

The split is not cosmetic. Reading the highest-traffic engines most often, and the rest on a slower cadence, is what lets us cover nine engines at all rather than four. When you read a Stellarcast number, it comes with the engine attached, because a blended score across engines you do not care about is one of the easiest ways to be misled.

Honesty about what each engine actually is

Not every "AI engine" answers the same way, and flattening that difference is a quiet form of lying. Some engines answer from the live web and return citations you can inspect. Others answer from training knowledge alone, with no live retrieval and no sources. DeepSeek, for example, answers from training knowledge with no live web and no citations, so we label it exactly that way rather than implying it browsed the internet to reach its answer. Google AI Overviews, ChatGPT search, Gemini and Perplexity behave more like live, citing systems; Copilot draws on a search index behind it. We keep those distinctions visible in the data, because "you were not mentioned" means something very different on a live-web engine than on one answering from a year-old training snapshot.

"A number without a spread hides the noise, it does not remove it. So we report a band, not a single figure."

Want to see this run against your brand? Get a free visibility audit and we will show you the real per-engine picture.

The prompts: a fixed, native, per-locale panel

Everything starts with the questions. We build a query universe from the real, buyer-style questions people ask in a category, expanded across the topics a brand wants to watch, and we keep the panel fixed so that runs stay comparable over time. Change the questions and you change the measurement, so the panel is treated as a controlled instrument, not a thing to tweak between runs.

Crucially, prompts are native to each locale, not machine-translated from one base market. A German buyer's question is written in German the way a German buyer would phrase it, because a translated prompt measures a translated question, not the real one. Locale is a first-class part of every measurement we take, not an afterthought bolted on at the end.

What we read from each answer

Once an answer comes back, we extract a small number of things that actually matter, and we compute them per engine and per locale rather than as one global blur.

Rivals are not guessed by string-matching a name. Competitors are confirmed entities, so a "rival cited while you are absent" signal means a real, named competitor, not a coincidence of wording.

Why we report a band, never a bare number

This is the rule we are most stubborn about. AI answers are non-deterministic: ask the same engine the same question twice and you can get two different answers. A single "your visibility is 34 percent" would imply a precision that does not exist. So we do not show a bare number. We show a range across repeated runs, so you can see the spread, a percentile band rather than a single point. A tight band means a stable result you can act on; a wide one means the engine is genuinely inconsistent about you, which is itself something worth knowing. Hiding that behind one clean figure would be the most flattering thing we could do and the least honest.

The cadence: measurement is a signal, not a project

Citation patterns move, sometimes sharply, within weeks. A source that dominates an engine's answers this month can be a minor player the next. That is why we re-measure on a schedule rather than running a one-off audit and calling it done. Core engines are re-read on the fastest cadence, extended engines less often, and every reading is stored as an append-only, date-stamped time series with the engine's model version and the prompt version attached to each row. Nothing is silently overwritten, so when a number moves, we can trace exactly what changed and when. A measurement without a date, in this space, is barely a measurement at all.

Proving lift honestly, including when we cannot

The hardest discipline is the last one: saying whether a change actually worked. When a brand makes a fix, we compare visibility before and after, but we refuse to overclaim. Every lift we report carries a confidence label. A result is marked measured only when the evidence is strong enough, likely when it points that way but is not yet conclusive, and too-early when not enough time or data has passed to say anything at all. "Too early" is a real, common answer in our system, not a failure. Telling you a change worked before the data can support it would be exactly the kind of confident-but-wrong claim this whole methodology exists to prevent.

The rules, in one place

If you take nothing else from this, take the short list of commitments the measurement holds itself to. They are the same tests we suggested you apply to any tool, including ours.

The takeaway

Good measurement in AI search is mostly a set of refusals: refusing to show a single number when the truth is a range, refusing to blend engines that behave nothing alike, refusing to translate a prompt and call it the local question, and refusing to claim a lift before the data earns it. That is the discipline behind every figure we publish, and it is why we are comfortable being this open about the method. If a number is worth showing you, it is worth showing you how it was made.

See your real per-engine picture

This is the method behind every Stellarcast audit: nine engines, every locale, ranges not guesses, and citations classified as yours, earned or a rival's. Request a free visibility audit and see where you actually stand.

Get your free visibility audit

Frequently asked questions

Which AI engines does Stellarcast measure?

Nine. We split them into core engines, measured on the fastest cadence because they carry the most traffic, and extended engines, measured less often. The core set is Google AI Overviews, ChatGPT, Gemini and Perplexity. The extended set adds Google AI Mode, Claude, Copilot, Grok and DeepSeek. We label each engine honestly: some answer from the live web with citations, while others, such as DeepSeek, answer from training knowledge with no live web, and we mark them as such rather than pretending every engine works the same way.

Why does Stellarcast show a range instead of a single visibility number?

Because AI answers are non-deterministic: the same question does not reliably return the same answer. A single clean number would be false precision. Instead we report a band, showing the spread across repeated runs rather than one figure, so you can see how stable or noisy a result really is. A number without a spread hides the noise, it does not remove it.

How does Stellarcast measure share of voice?

We run a fixed set of real, buyer-style prompts against each engine on a schedule, then read the answers for two things: whether your brand is mentioned, which gives a visibility rate and a mention-share of voice against your rivals, and which pages are cited, which gives a citation share of voice. Every result is computed per locale and per engine, date-stamped, and rolled up over a time window so the trend, not any single run, is what you read.

How are cited sources classified?

Every citation an engine makes is resolved to a URL and domain and classified as owned, meaning your own properties, earned, meaning third-party pages that mention you, or a rival's. That is what lets us tell you not just whether you appear, but whether the answer is being built from your own material, from independent corroboration, or from a competitor's page instead.

How often does Stellarcast re-measure?

Continuously, on a schedule, because citation behaviour shifts within weeks. Core engines are re-measured on the fastest cadence and extended engines less frequently. Every reading is stored as an append-only, date-stamped time series with the engine model and prompt version attached, so a change in the numbers can always be traced and nothing is silently overwritten.