Key points
- Measurement answers several questions, not one: is the brand named, in what context, are the facts correct, does the model cite your sources, and where does it send the purchase.
- Cadence: monthly while you are actively working on data and content, quarterly once catalogue and brand are stable. Anything less frequent breaks the link between a change in results and its cause.
- Model answers are non-deterministic, so ask every question several times and record frequency. Draw conclusions from a trend across measurements, not from a single screenshot.
- Fix in this order: wrong facts, then missing data, then content, then external signals. The order follows from what does the most damage and what is fastest to repair.
The external signals from the previous chapter share one inconvenient property: they work slowly and largely outside your control. The same is true of everything that came before. Product data, content and mentions only start working once models pick them up, and you will only learn that they did if you check. The audit in chapter 2 was a snapshot of a single day. This chapter turns it into a process: what to measure, how often, how to record results so they are still comparable six months from now, and in what order to fix what the measurement shows.
What to measure: several questions, not one number
The temptation is to reduce AI visibility to one percentage on a slide. I would advise against it, because a single percentage never tells you what to fix. A workable measurement answers several questions, asked separately for each query in your control set.
| Control question | What you check | What you record |
|---|---|---|
| Is the brand named at all? | Whether the answer contains your brand or product name | Yes or no, separately per model |
| In what context? | Whether you are recommended, mentioned in passing, or cited as the weaker option | A short quote of the passage that concerns you |
| Are the facts correct? | Price, availability, category, product attributes, markets served, company status | A list of errors and, where visible, the source the model took them from |
| Does the model cite your sources? | Whether your pages appear among the links given, or only third-party sites | A list of cited domains |
| Where does the purchase go? | Whether the product is attributed to your store, a marketplace or a competitor | The purchase destination named in the answer |
Five columns in a spreadsheet are enough to start. What matters more than the tool is that the definitions stay identical at every subsequent measurement.
The naming is secondary, but worth knowing because it shows up in tool pitches. The industry uses terms such as Share of Model for a brand’s share of answers across a fixed set of queries. In our own service we work with three: Brand Visibility (is the brand named), Product Mention Rate (does the specific product appear when the question fits it) and Store Attribution (is the purchase credited to your store). I deliberately give no reference values here, because meaningful benchmarks do not exist: results depend on category, language, the composition of the query set and the model. The only honest comparison is against your own previous measurement.
What has to stay fixed
The query set you built during the audit now becomes an asset, and its greatest value is that it does not change. Four things must stay constant at every measurement: the wording of the questions, word for word, the language and market, the list of models and the scoring method. Add new questions as a separate group and compare them only from their own first measurement, rather than mixing them into the old pool and breaking comparability across the board.
The conditions under which you ask are part of the method too, though they are rarely written down. Ask in a session without chat history and without personalisation, ideally logged out, otherwise you will measure your own earlier footprints instead of what a stranger sees. Also note whether the model used web search for a given answer, because an answer with search and one without are effectively two different measurements and should not be averaged together.
How often to measure
Monthly, if you are actively working on data, content and external signals, because you want to see the effect of your own changes. Quarterly, if catalogue and brand are stable and the measurement mainly guards against things quietly breaking. Less often than quarterly stops being useful: across that gap you can no longer connect a change in results to any specific cause.
More often than monthly is usually pointless, for a mundane reason: a change on your site has to be fetched, processed and settle into the sources a model draws on, and that takes time. Outside the regular cadence, it is worth measuring in four situations: after a platform migration, after a name or domain change, after a major catalogue rebuild, and after a widely reported change on the model provider’s side.
How to record results
The format has one purpose: to let you reconstruct, six months later, exactly what you saw. The minimum is the measurement date, the model name and version where visible, market and language, the exact question, the full answer copied in its entirety, the scoring against the questions in the table above, and the list of sources the model provided.
Two mistakes come up most often. The first is recording the score without the raw answer; a quarter later nobody can reconstruct why the context was judged neutral at the time, and arguing about it costs more than the measurement itself. The second is collecting screenshots instead of text, which makes results impossible to search or count. A plain spreadsheet with one answer per row holds up far longer than people expect, and it ports into any tool you adopt later.
Why a single measurement means nothing
Model answers are non-deterministic: the same question asked twice in a row can produce a different answer, a different set of recommended brands and a different set of sources. This comes from how text is generated, from queries being routed to different model variants, from web search being used or skipped, and from ordinary updates to the underlying sources. According to industry sources, comparisons of answers to identical questions show the divergence extends beyond phrasing to who gets named and what gets cited.
The methodological conclusion is simple: ask every question several times and record frequency, not the bare fact that your brand appeared once. A single appearance is not a success and a single absence is not a failure. The signal only emerges from repeatability within one measurement and from the direction of change across the next ones.
Rule of thumb: one answer is an anecdote, not a measurement. Base conclusions on three consecutive measurements with the same query set, not on a screenshot somebody forwarded on a Friday afternoon.
What to do with the results: order of repairs
A measurement usually produces a long list of things to do, and this is exactly where teams stall. The order we use follows a simple calculation: first what does the most damage and is fastest to repair.
1. Wrong facts. The model quotes an outdated price, claims you do not ship to a given country, confuses you with another company or describes a pre-rebrand offer. Fix this immediately, because that kind of error actively costs sales, and its source can usually be identified and corrected either on your side or on one specific external site.
2. Missing data. The model has nothing to work with: attributes are absent, the feed is incomplete, structured data covers only part of the catalogue. This is the chapter 3 work, expensive once but long-lasting and applied across the whole catalogue at once.
3. Content. The data exists, but nobody answered the question a customer actually asks. This is chapter 4, meaning the most editorial effort and an effect spread over time, though also the most durable.
4. External signals. Reviews, rankings and mentions from chapter 5. The slowest layer, so start it in parallel with the rest, but do not expect it to move the next monthly measurement.
On your own or with a partner
Doing this yourself makes sense more often than tool vendors suggest. One market, one language, a few dozen control questions and a person on the team who will genuinely keep the cadence: that is a spreadsheet and roughly an hour a month. The biggest risk is not the absence of a tool but the fact that after the third month nobody runs it any more, at which point you lose your reference point entirely.
A partner or a tool starts paying off with several markets and languages, a large catalogue, several models to check, and when the measurement is supposed to produce concrete repairs in data and on the site rather than a chart for a meeting. That is how we run it at Seedlight as our AI Visibility (GEO) service: a fixed query set, per-model measurement, Brand Visibility, Product Mention Rate and Store Attribution, plus the fixes in the data and content layer. The caveat belongs here, stated plainly: nobody guarantees a position, a mention or a citation, because those depend on variables no provider controls.
The whole path in one paragraph
We started with where models get their knowledge about your store and why your own site is only one source among many (chapter 1). Then came the audit, checking what AI says about you today and writing down the baseline (chapter 2). On that foundation you put product data in order, feed and structured data as a single source of truth about products (chapter 3). Next you wrote content an answer can be lifted from, instead of content that merely reads well (chapter 4). Chapter 5 took the work beyond your own domain, into reviews, rankings and mentions that corroborate what you say about yourself. This chapter ties it into a cadence: measure, record, repair in a set order, measure again.
One thing has to be said plainly at the end, because without it the whole guide would be dishonest. You do not control what a model answers. Providers change models and sources without notice, the way citations are selected stays opaque, and nobody, ourselves included, can guarantee you a mention or a citation. What you do control is the accuracy of facts about you, the completeness and accessibility of your data, the quality of your answers to real customer questions, the consistency of your brand across the web and the regularity of your measurement. That list is long enough to fill a year of work and concrete enough to start tomorrow. The rest, including the fact that some customers will decide without ever visiting your site, is a consequence we cover separately in our piece on zero-click commerce.
Questions
How often should I measure visibility in AI answers?
Monthly if you are working on product data, content and external signals in parallel, because you want to see the effect of those changes. Quarterly if catalogue and brand are stable. Less often than quarterly loses diagnostic value, because a change in results can no longer be tied to a specific cause.
Why do I get a different answer every time I ask the same question?
Because model answers are non-deterministic. The same prompt may be routed to a different model variant, may or may not trigger a web search, and text generation involves randomness by design. That is why a single result is not a measurement: ask each question several times, record how often your brand appears, and read the trend across several consecutive measurements.
Do I need a paid tool to measure AI visibility?
Not to begin with. For one market and a few dozen control questions, a spreadsheet and an hour a month are entirely sufficient, provided the definitions and the query set stay unchanged. A tool or a partner starts paying off with several languages and markets, a large catalogue, and when the measurement is meant to drive concrete repairs rather than produce a report.