BEAM

Seedlight BEAM: our framework for launching, automating and growing eCommerce platforms →

← All articles
AutomationsSzymon Żynda8 min read

AI catalog translation without an agency: where translation ends and localization begins

Bulk catalog translation is one of the best first AI projects in a store: it runs on data you already own, and you catch every mistake before it goes live. On one condition, that you know where translation ends and localization begins, because that is where most failed rollouts break.

Bulk catalog translation is one of the best first uses of AI in a store: the work repeats, the source material already sits in your systems, and you review the output before anything reaches a customer. There is one condition, and it is where most failed rollouts break: translation is not localization. A model will carry sentences from one language into another, and it will do it fast. It will not convert a size chart, it will not name a category the way the target market names it, and it will not tell you whether your marketing claim is even allowed there. That boundary decides whether the project ends with a catalog ready to sell or with correct-sounding text that does not convert.

Key takeaways

  • Catalog translation is a strong first AI project: the work repeats, the source material is already yours, you can judge quality in minutes, and a bad result gets replaced before any customer sees it.
  • Translation is not localization. Size charts, category names, mandatory product information and marketing claims stay with a human or with a rule in your data, not with a language model.
  • The glossary decides whether the whole project holds together. Without one, the same term gets a different translation in every batch, and you hear about it from a customer.
  • Your reviewer has to know the target market, not just the language. The test is not "is this understandable", it is "does this sound like a store from that market and do the specs hold".

Why catalog translation makes a good first AI project

Before you pick a tool, answer a different question: is this task even suited to being your first one. Catalog translation is, for four reasons that rarely line up together:

  • The work repeats: a thousand descriptions is the same task a thousand times, not a thousand different problems. That is exactly the shape of work worth automating.
  • The data already exists: nothing has to be collected first. The source material sits in your catalog, in fields your team has been filling in for years.
  • Quality is quick to judge: reading a dozen product pages tells you whether the output is acceptable. You are not waiting a quarter for sales data to answer the question.
  • A bad result costs you nothing but time: as long as the translation is unpublished, a mistake costs an afternoon, not an order.

That last point matters more than it looks. Plenty of first AI projects collapse because the team picked a task that could not be undone or could not be checked quickly. Catalog translation has the opposite risk profile: the control gate sits before publication, so mistakes stop on your screen. We apply the same qualification test more broadly to operational work in the piece on the anatomy of automation savings.

The prerequisite: a catalog that can be translated at all

A model translates whatever it is handed. If specifications live inside your description text instead of in fields, they will still live inside the description after translation, only in a second language. The mess does not get cleaned up along the way. It gets copied to the next market, and from then on you maintain it in two places instead of one.

CATALOG CSV · MESSY EXPORT TO CLEAN ROWSRAW EXPORTS,M,L in one cellEAN checksum?42cm / 0,42 mbroken encodingquotes, separatorsduplicates, gapsCLAUDE CODEsplit variantsvalidate GTINnormalize unitsfix encodinggenerate slugsflag dupes + gapsworks on the file, not a chatCLEAN CATALOGone row per variantvalid check digitsone unit formatclean UTF-8unique slugsflagged for reviewwork on a copy · review the diff before you save · no secrets in the file

A typical input looks like this: the description reads "100% linen, 42 cm long, ecru", and the attribute fields are empty because nobody ever filled them in. After translation you have the same sentence in German and still no data for filters, for the channel feed or for a comparison engine. The translation is correct and the catalog is still not ready to sell.

So the order of work is the opposite of what impatience suggests: clean the catalog first, translate second. Splitting specs out of prose into fields, normalizing units and mapping colors to a controlled list is its own task, and AI handles it well: we walk through cleaning a catalog CSV step by step separately. If the source-language descriptions do not exist yet, that is a third task, closer to generating product descriptions than to translating them.

The rule: translation copies your catalog structure onto a new market, good or bad. If specifications currently live in sentences instead of fields, translation is not the first job on the list, it is the second one. A language model will not fix your data structure, because that is not a language problem.

The glossary is the heart of the project, not an accessory

A glossary is a list of terms with a decision attached to each one: what happens to it in the target language. It sounds like paperwork, and it is the only thing holding a catalog together when it is translated in batches across many days. What belongs in it:

  • Proper nouns and brands: what stays untouched and what has a local form. A translated model name that customers then cannot find in search is a classic own goal.
  • Product and range names: one decision for the whole family, so the same range is not called three different things on three product pages.
  • Industry terminology: technical vocabulary that has one settled form in the target language, even when a dictionary offers three equally correct ones.
  • Prohibited claims: phrasing you may not use in that country or in that category. The model does not know your list until you hand it over.
  • Units and formats: number notation, decimal separator, date format, and whether units get converted or stay as they are.

Without a glossary the same term picks up three different translations across three batches, and you will not find out until a customer asks how two products differ when they are in fact the same one under two names. The glossary also grows during the project: every reviewer correction that follows a pattern should come back as a rule, instead of ending its life as a one-off edit in a single file.

Where translation ends and localization begins

Translation moves sentences between languages. Localization makes the offer correct and credible in one specific market. A model does the first job very well and the second only in appearance, because it has no access to knowledge that is not in the text. Four areas stay with humans and with rules in your data:

  • Size charts and units: this is conversion, not translation. It belongs to a conversion table in product data, not to a language model.
  • Category naming: the target market searches for its own word, not for a literal rendering of yours. Trainers and sneakers describe the same shoe to two different audiences.
  • Mandatory information: the wording required by law, warnings and data that must appear on a product page in that country. Here the settled phrasing wins over the elegant one.
  • Brand voice and marketing claims: a promise that is industry standard in one country can be off limits in another. That call belongs to someone who knows that market.

The split of responsibility is clearest element by element on a product page. The further down the table, the less work is left for the model and the more sits with your data and your judgement:

Element of the product pageWhat the model doesWhat stays with a human
Marketing descriptionTranslates, adjusts length and structureBrand voice and whether the promise is allowed in that market
Attributes and specsCarries values across from fields, not from proseField completeness and units matching the local standard
Size chartsNothing useful. A translated number is the same numberA conversion rule in the data plus a size table per market
Category namesOffers a literal equivalentChoosing the word the target market actually uses
Legally required contentNothing, until it is given the approved wordingSupplying that wording and checking it is complete
Transactional messagesA draft at most, if you let it near them at allSigning off every sentence a buyer sees after checkout

The model owns the language layer. Your data and your people own correctness in the market.

A size 38 is still a 38 after translation

Sizing is the most instructive example of this boundary. Sizing systems differ between markets, so a number that means one thing in your home country means something else, or nothing at all, somewhere else. The model will copy it across unchanged, because from its point of view there is nothing to translate, and it will be formally right. You will not see the consequence in a translation report. You will see it in the returns report.

The practical conclusion: derive sizes and units from data, not from text. A conversion table per market, applied as a rule when the product page is generated, is cheap, repeatable and testable. A language model is the most expensive and least predictable tool you could point at that particular job.

A workflow with a gate: this is a loop, not a one-off job

The most common organizational mistake is treating catalog translation as one enormous job: you push six thousand items through, you get six thousand translations back, and you are left with a pile nobody will ever review. Work in batches instead, in a loop where each batch comes out better than the last:

  • The batch: the model translates a slice of the catalog, one category for instance, with the glossary and the formatting rules in context.
  • The sample: the reviewer does not read everything, only a sample from the batch, picked so that the hard cases are represented in it.
  • Corrections into the glossary: every fix that follows a pattern returns to the glossary as a rule, instead of staying in one file.
  • The next batch: starts with the extended glossary, so the same mistake does not repeat across the next few thousand items.
APPROVED WITHOUT EDITS · BY BATCH54%B163%B271%B378%B483%B587%B6typical tuning curve: reviewer corrections feed back into the rules

The output of that loop is a rising share of translations accepted without edits. The chart above is illustrative: it shows the shape of the curve, not the result of any particular rollout. Your numbers will depend on catalog quality, on how precise the glossary is, and on how hard the category is linguistically. Watch the direction rather than the value: if the share stops rising after a few batches, corrections are not making it back into the rules and your reviewer is fixing the same thing over and over.

The gate exists so that nothing is taken on trust. Nothing reaches the storefront without human approval. Once the unedited acceptance rate settles high, you shrink the sample rather than remove the gate. Publishing with no check at all is not the mature version of this process, it is the absence of one.

Who reviews, and against which test

This is where most projects quietly lower the bar. A reviewer who knows the target language from school will judge whether the text is understandable. A reviewer who knows the target market will judge whether it is credible, and that is an entirely different question. The gap is invisible in the text and visible in how buyers behave.

So state the test in hard terms. It is not "is this understandable", it is: does this sound like a store from that market, and do all the specs hold. The first half is market judgement, the second half is a data check. One person has to be able to do both, or the two roles get split between two people and both sit inside the gate.

In practice this means answering one question before you start: who is going to review this. If the honest answer is "nobody", that is not a problem a better prompt will solve. It is the binding constraint on the whole project, and it has to be removed before the first batch runs.

What not to automate

Not everything in a store is text you may generate and then spot-check. Keep three categories outside the automation, or behind full approval rather than sampling:

  • Legal text: terms and conditions, the returns policy, warranty and withdrawal information. Settled wording is the point, and a mistake has formal consequences, not stylistic ones.
  • High-stakes transactional messages: order confirmations, payment, refund and complaint notices. Customers read these at the moment when they forgive the least.
  • Content where an error is expensive and hard to spot: usage instructions, warnings, safety parameters. If a mistake would not show up in a sample, sampling is the wrong control.

When you do want an agency

The honest answer: when you are entering a market nobody in your company knows, and you have no one to put on review. AI does not remove the need for market competence, it moves that competence from writing to judging. If it does not exist in the company, it does not exist in the AI process either, and you end up producing content quickly that nobody is able to evaluate.

The sensible middle option: an agency or a native speaker handles the first batch and builds the glossary and tone rules with you, then AI scales the rest of the catalog on those rules, with periodic spot checks from the same people. You are buying judgement and a standard, not billable hours spent retyping six thousand descriptions. Entering a market you do not know, that is usually the cheapest route to quality you can sustain.

AI catalog translation is not a choice between an agency and a machine. It is a choice about what you hand to the model (language, at scale, inside a loop with a gate) and what stays a rule in your data or a decision by someone who knows the market. The wider context, including the order in which AI is worth introducing into store operations, sits in our guide to AI for eCommerce. What this looks like wired permanently into a platform is described on the AI catalog translations page.

FAQ

Is AI enough to translate a store for a new market?

For the language layer, usually yes, on two conditions: a clean catalog and review before publication. It is not enough for localization, meaning size charts, category names, mandatory information and the judgement of whether a given marketing claim is permitted in that market.

What is the difference between translation and catalog localization?

Translation moves sentences between languages. Localization makes the offer correct and credible in a specific market: it converts sizes and units, uses the category names local buyers actually search for, and respects the information required on a product page there.

Who should review AI catalog translations?

Someone who knows the target market, not only the language. The test is not "is this understandable" but "does this sound like a store from that market and do the specs hold". Having no such person is the binding constraint on the project, not a prompting problem.

What should never be translated automatically?

Legal text, high-stakes transactional messages, and content where an error is expensive and hard to spot, such as usage instructions, warnings and safety parameters. Those go through full human approval rather than sampling.

Journal

Szymon Żynda

Co-founder of Seedlight · eCommerce platforms, AI, SEO and GEO

More by this author

Newsletter

The Journal, straight to your inbox

New articles and lessons from real builds, every now and then. No spam, unsubscribe with one click.