BEAM

Seedlight BEAM: one place to run your whole eCommerce, with AI agents that know your business →

← All articles
eCommerceSzymon Żynda9 min read

AI product recommendations: how much data you need before they earn anything

Recommendations learn from interactions per item, not from total orders. A calculation you can run on your own numbers, four approaches with the data threshold for each, and why measuring without a control group always reports success.

AI product recommendations earn their keep when each item in your catalogue carries enough interactions for a model to draw conclusions from. Revenue size does not decide that, the ratio of behaviour to products does. A store with two thousand orders and three hundred items has dense data. A store with the same orders and twenty thousand items has data thinned to nothing, and it will get recommendations that look random, because they are.

Key takeaways

  • What matters is interactions per item, not your total order count. Two stores with identical revenue can have working recommendations and completely dead ones if their catalogue sizes differ.
  • The calculation is simple: monthly orders times items per order, divided by active SKUs. Below roughly one interaction per item per month, behavioural learning has nothing to work with.
  • At a low data threshold, attribute-based recommendations beat behavioural learning, because they work from day one and need no history. In exchange they need a tidy catalogue.
  • Without a control group every recommendation report looks good, because it credits sales that would have happened anyway. The only honest measure compares against a randomly switched-off slice of traffic.

Below is a calculation you can run on your own numbers in five minutes.

Comparison of two stores with identical monthly order counts: the first, with a catalogue of three hundred items, reaches around twenty interactions per item and qualifies for behavioural learning, while the second, with twenty thousand items, reaches three hundredths of an interaction per item and qualifies only for recommendations built on product attributes.
Same revenue, two different verdicts. Data density follows catalogue size, not order count.

Recommendations learn from interactions, not revenue

A behavioural engine looks for patterns in what people viewed and bought together. To establish that product A goes with product B, it has to see that pair repeatedly across different people. With three hundred items and a few thousand events a month those repetitions are plentiful. With twenty thousand items the same events spread so thin that most pairs appear once or never, and then the model learns noise and serves it back as a pattern.

Vendors call this the cold start problem and usually say it passes. With a large, slowly rotating catalogue it never passes.

Simple arithmetic on your own numbers

Take your monthly orders, multiply by average items per order, divide by active SKUs. The result approximates purchase interactions per item per month. Product page views raise that figure roughly twentyfold, but they carry a far weaker signal than a purchase, so count conservatively and treat the result as a floor.

Store profileOrders/monthActive SKUsInteractions per itemVerdict
Narrow D2C brand2,000300~13Behavioural learning works
Mid-size specialist store2,0003,000~1.3Borderline, needs attribute support
Broad technical catalogue2,00020,000~0.2Attributes and rules only
Large general merchant50,00020,000~5Behavioural learning works
B2B wholesale, repeat orders8005,000~0.5Account history rather than a model

Our own calculation, assuming two items per order. Thresholds are indicative and depend on catalogue rotation. Treat this as a template to calculate with, not as study findings.

Look at rows one and three. Same revenue, and the verdict differs purely because of catalogue size.

Four approaches and the data threshold for each

The phrase „AI recommendations" covers techniques with wildly different requirements. Choosing between them is really a decision about which asset you hold: behavioural history or a tidy catalogue.

ApproachWhat it needsWhere it fitsWeak point
Manual rulesCategory knowledge, no dataLaunch, small catalogue, accessoriesDoes not scale, ages silently
Attribute similarityComplete product attributesLarge catalogue, low trafficSuggests similar items, rarely complementary ones
Behavioural learningDense history per itemSmall catalogue or heavy trafficCold start on every new item
Language model over the catalogueDescriptions and attributes in one recordDescriptive search, adviceQuery cost, risk of invented pairings

Our own comparison. In practice, stores above the threshold combine the attribute approach with behavioural learning so new stock does not wait for data.

The attribute approach is the most underrated row here. It needs zero order history, works from day one, and asks for exactly what an AI assistant in your store asks for: a complete product record with unit, variant and attributes in separate fields. The same catalogue work serves both, which makes it the best returning investment in this area.

Fast catalogue rotation wipes out the learning

Data density is one variable, item lifetime is the other. When half your catalogue turns over each season, the model learns products that will be gone in three months, while new arrivals enter with no history exactly when their margin is highest. Fashion and home furnishing have this problem structurally. Spare parts and technical chemicals do not have it at all, because the same item sells for years and history accumulates.

That is why a high-traffic fashion store can be a worse candidate than a wholesaler with a tenth of the traffic.

Measuring without a control group always reports success

The most common post-launch report reads: „the widget generated 8 percent of revenue". That number says nothing, because it credits recommendations with every sale where a customer clicked a tile, including products they would have bought anyway and ones they would have found through your search box. An honest measure requires switching recommendations off for a random slice of traffic and comparing revenue per session between groups. The gap between groups is your result, and everything above it is attribution.

  • Switch recommendations off for a random 10 percent of sessions for at least two full weeks.
  • Compare revenue per session and basket value rather than clicks on the widget.
  • Check returns in both groups: a recommendation that lifts both sales and returns has earned nothing.
  • Repeat after a season, because data density and catalogue mix shift over time.

Three stores where recommendations will never pay back

  • One-off purchase catalogues. Made-to-measure furniture, equipment bought once a decade, services. There is nothing to recommend alongside, and one customer’s history says nothing about the next.
  • Stores with a few dozen items. Your customer will see the whole assortment anyway. Good navigation and a hand-picked related list do the same job for a fraction of the cost.
  • Catalogues with neither attributes nor traffic. Here the missing piece is not an engine, it is both assets at once. Money goes into product data first, which we covered in our piece on what an AI agent sees in your store.

What we recommend. Work out your interactions per item before asking anyone to quote an engine. Below one interaction a month, pick attribute-based recommendations and spend the budget on tidying the catalogue, because that same work also serves search, feeds and visibility inside model answers. Above five interactions, behavioural learning makes sense and a deployment conversation is worth having. Wider context sits in our guide to AI for eCommerce, and conversion plus product data work belongs to the Maintenance & Growth stage.

FAQ

How many monthly orders do I need before AI recommendations make sense?

Order count alone does not settle it. Divide monthly orders multiplied by items per order by your active SKU count. Above roughly five interactions per item, behavioural learning has something to work with; between one and five it needs attribute support; below one, only recommendations built on product features function at all. A store with 800 orders and a narrow catalogue can be a better candidate than one with 5,000 orders and twenty thousand items.

Do language model recommendations solve the cold start problem?

Partly, and that is their strongest suit. A language model reads the description and attributes, so it can propose something sensible for stock that arrived yesterday with no history. The price is query cost at high traffic and the risk that the model pairs products that do not belong together because it found surface similarity in the copy. In practice it works better as a supplement for new stock than as your only engine.

How does personalisation differ from product recommendations?

A recommendation answers what to show next to this product. Personalisation changes what a given person sees across the whole store: listing order, homepage content, which collections appear. Personalisation requires recognising a returning visitor and usually consent to process their data, so its legal and technical entry bar is higher. At low traffic it also delivers less, because most sessions are first visits the system knows nothing about.

Can I measure recommendations without a control group?

You can measure them, but the result will not tell you what they added. Without a control group you credit recommendations with sales that would have happened regardless, so every report comes out positive. The minimum honest version is switching the widget off for a random ten percent of sessions for two weeks and comparing revenue per session. If your vendor discourages that test, that is information in itself.

Journal

Szymon Żynda

Co-founder of Seedlight · eCommerce platforms, AI, SEO and GEO

More by this author

Newsletter

The Journal, straight to your inbox

New articles and lessons from real builds, every now and then. No spam, unsubscribe with one click.