BEAM

Seedlight BEAM: our framework for launching, automating and growing eCommerce platforms →

← AI visibility for ecommerce

Chapter 2 of 6

The audit: what AI says about you today

A repeatable afternoon procedure: which questions to ask assistants, how many assistants to use, how many times to repeat, and what to record in a spreadsheet so you end up with a baseline before you change anything. The headline number of a first AI visibility audit is not how often you are mentioned, it is how many facts about you are wrong.

8 min read

Key points

  • The main output of a first audit is a list of wrong facts about you, not a mention count. A wrong fact actively takes customers away from you, and it is the fastest thing to fix because it has a traceable source.
  • Ask 10 to 15 questions across four types (brand, category, comparison, problem), on at least three assistants, three times each, in fresh chats with no history or personalisation.
  • One spreadsheet row is one answer, not one question: date, assistant, repetition, whether the brand appeared, in what context, which facts came up, whether they are true, which sources the model cited.
  • Freeze the result with a date as your baseline before you change anything. Model answers are non-deterministic, so only a trend across a fixed question set is comparable, never a single answer.

Now that you know where a model gets its knowledge of your store, the next question asks itself: what does it actually say about you today? Most teams know the answer only from anecdotes, because somebody once asked ChatGPT about the brand and either felt flattered or offended. That is not enough to build on. This chapter is a procedure you can run in a single afternoon, and it gives you a reference point before you start changing things. Do it before any fixes, because without a starting point you cannot tell your own effect from coincidence.

The rule that frames the whole audit

The first instinct is always the same: count how many times we get mentioned. I would not start there. At the outset, the number that matters is how many facts about your store are wrong, not how often you are named. The reason is practical. A missing mention is a lost opportunity, unpleasant but passive. A wrong fact works against you actively: if the model tells a customer you do not ship to Czechia, that your returns window is fourteen days instead of a hundred, or that a product is not in your range, you have just lost that sale and you will never hear about it.

There is a second, less obvious reason. Mention counts fluctuate between repetitions of the same question, so on a small sample they are simply unstable. The list of wrong facts behaves differently: it is remarkably stable, because a model repeats the same distortion until you fix the source. That makes it measurable and, more importantly, fixable. So treat the first audit as a technical inspection rather than a popularity contest.

Output number one from your first audit is a single page: "here is what models say about us that is untrue, sorted by how much it costs us". Price, delivery cost and coverage, return terms, availability, company status. That list is the one thing from the whole audit worth showing the board at the first meeting, because every line translates into lost orders and every line has a specific owner on your side.

Four types of test queries

Build the question set from what customers ask, not from the phrases in your SEO sheet. You are testing a buying channel, not your own vocabulary. You need 10 to 15 questions across four types, a few of each:

  • Brand: "what kind of store is [name]", "is [name] a trustworthy store", "what are delivery costs and return terms at [name]". You are checking which facts the model holds about you and how many of them are true. This is where the most important list of the audit comes from.
  • Category: "where to buy [category] in Poland", "recommended stores for [category]", "a good [category] store with fast delivery". You are checking whether you make the shortlist the customer actually chooses from.
  • Comparison: "[your brand] or [competitor]", "where is [product] cheaper", "which [category] store has the best return policy". You are checking how you look next to competitors and which criteria the model compares you on. Those criteria are worth more than the verdict.
  • Problem: "how to choose [product] for [use case]", "what to buy for [occasion] under [budget]", "how does [variant A] differ from [variant B]". The customer does not know your brand yet and asks about their problem. This is where guide-style content usually wins, not product pages.

If you sell in several markets, ask the questions in your customers' languages and run the set separately per market. This is not a formality. In practice the brand picture is often accurate in Polish and noticeably thin in Czech or German, simply because the model has less material about you there. A single blended "AI visibility" score would hide that gap, and it is usually the cheapest one to close.

How many assistants, how many repetitions

The minimum is three different systems, for example ChatGPT, Perplexity and Gemini, plus a look at what AI Overviews shows in Google for your category questions. Ask each question three times, each time in a fresh chat. The reason is fundamental and will come up again in this guide: model answers are non-deterministic. The same question asked twice can produce different wording, a different order, and sometimes a different set of brands. A single answer is an anecdote; only repetition shows you what the rule is.

Run the test logged out, or with memory and personalisation switched off. Otherwise you are testing your own conversation history rather than what a customer hearing about you for the first time will see. Work out the scale up front: 12 questions across 3 assistants with 3 repetitions is 108 answers. That is a few hours for one person, ideally the same person throughout, so the "correct or wrong" judgment stays consistent. Split it into two sittings, but keep it within a few days, because models get updated and an audit stretched over weeks mixes two different states of the world.

What to record in the spreadsheet

Build the sheet in whatever tool you like, as long as everyone fills it in the same way. The key rule: one row is one answer, not one question. Without that you lose the spread, and the spread is what tells you how shaky your position really is.

ColumnWhat you recordWhy you need it
Date and assistantDay of the test, assistant name, and whether the answer came with a web search (citations present) or notWithout a date the result is useless a quarter from now; the answer mode explains most differences between assistants
Query and its typeThe exact wording of the question and its type: brand, category, comparison, problemYou count results separately per type, because gaps almost never spread evenly
Repetition number1, 2 or 3Shows the spread: named once in three attempts is a very different situation from named every time
Was the brand mentionedYes or no, with no interpretationBasic presence, later counted as a share of answers, never as a position
In what contextAs a recommendation, as a passing mention, as an example of something to avoidMentioned is not the same as recommended, and that difference is sometimes the entire problem
Who else is mentionedThe other stores and brands, in the order they appearThis is your real comparison set in this channel, often different from the one in your strategy deck
Facts about youEvery specific claim: prices, delivery cost and time, returns, range, years in business, countries servedThe raw material for the most important list in the audit
Fact assessmentCorrect, out of date, or wrongOut of date and wrong need different fixes: the first is a refresh problem, the second is often a conflicting-sources problem
Cited sourcesThe domains and URLs the model referred toTells you directly where the model reads about you; a ready-made list of places to fix
Copy of the answerPasted text or a screenshotA quarter from now you will not reproduce that answer, because models change; without a copy you are left with memory

Ten columns are enough for a first audit. More important than the tool is that the "correct, out of date, wrong" judgment follows the same rule throughout.

Turning the spreadsheet into a baseline

Pull three things out of the collected answers, and only three. First, the list of wrong and out-of-date facts, sorted by what they cost you. Second, the share of answers in which the brand appeared at all, counted separately per question type, because presence on brand questions says nothing about presence on category ones. Third, two supporting lists: the domains cited, and the brands the model names instead of you.

Then freeze it: date, who ran it, which question set, which assistants. From that moment you do not change the method, because every change to the question set breaks comparability with the next measurement. If you must change something anyway, record it in the sheet as a separate version of the set. One honest caveat to close on: this is not a laboratory measurement, and you will not derive a "position in ChatGPT" from it, because no such quantity exists. Do not promise yourself or your board that the model will start recommending you after the fixes. You control the facts the model relies on, not the answer itself.

If you want to look at this more broadly than a one-off inspection, we go deeper into brand visibility as a standalone marketing metric in our piece on how AI describes your brand. Here we stay with the starting audit, and this is where the diagnosis ends. You have a list of errors, a reference point, and a list of the sources the model reads about you. The next three chapters fix all of it in order of leverage: product data first, then content, then external signals. Chapter 6 shows how to turn this one-off audit into a repeatable measurement.

Questions

How many questions and repetitions are enough for a first AI visibility audit?

To start, 10 to 15 questions across four types (brand, category, comparison, problem), asked on at least three assistants, three times each in a fresh chat. That comes to roughly a hundred answers, a few hours of work for one person. Fewer repetitions give you a result you cannot separate from chance; extra questions are better added at the second measurement than allowed to drag out the first.

Why does the same answer look different the second time?

Because model answers are non-deterministic: the same question can produce different wording, a different order of brands, and sometimes a different set of them. The answer mode matters too, meaning whether the model fetched data from the web or replied from memory. That is why you log the repetition number, count a share rather than a position, and draw conclusions from the trend rather than a single answer.

The audit found wrong facts about my store. Where do I start fixing?

Start with the facts that cost sales fastest: price, delivery cost and coverage, return terms, availability. Use the cited-sources column to find where the model got them, and fix them on your own site first, then in those external places. If the model answered from memory with no citations, the fix will take longer to show, because it first has to settle into the sources models read.

All chapters in this guide

AI visibility for ecommerce

  1. 01Where AI learns about your storeAn AI assistant does not look at your store, it reads the facts it can extract from it without ambiguity. This chapter explains how AI decides what to recommend and the four groups of sources behind that picture: product data, machine-readable content, external signals, and what the model remembers versus what it fetches live.
  2. 02 · You are hereThe audit: what AI says about you today
  3. 03Product data: feed and schema as the foundationVisibility in AI answers starts with product data, not with the blog. How to set identifiers, a complete Offer and shopping feed hygiene, so an assistant has a price, an availability status and a product identity to work with.
  4. 04Content that lands in answersProduct data gives a model price and availability, but the sentences about who a product is for and how it differs from the alternative have to come from your content. How to write so those sentences can be lifted off the page and used as they are.
  5. 05External signals: reviews, rankings and mentionsA model does not build its recommendation from your site alone. Reviews, industry rankings, media mentions and your marketplace listings act as verification of what you say about yourself. What you can build there, how fast, and what money cannot buy.
  6. 06Measurement and maintenanceHow to turn a one-off audit into a repeatable process: what to measure, how often, how to record results so they stay comparable, and in what order to fix what the measurement reveals. Plus an honest closing of the whole six-chapter path.