Key points
- An audit starts with a list of tasks and real hours, not a list of tools. A tool without a work inventory is a solution looking for a problem.
- Only the rule-based share of a task is a candidate for automation. Subtract residual oversight from it, and only that difference is a potential saving.
- Time is one criterion. The second is error cost and sales impact. Those set the order of implementations, not whether a task is worth doing at all.
- Record a baseline before you start: hours, error counts, time to publish, scale, and the date of measurement. Without it you cannot later prove the effect to yourself or to your board.
Now that the previous chapter has shown you where AI actually helps, the next question is which of it applies to your company. The map of areas is shared across the industry, but the order of implementation is specific to you and follows from numbers nobody else holds. This chapter is a procedure you run yourself, on paper or in a spreadsheet, before you talk to any vendor. It answers how to implement AI in a company in a way you can account for afterwards: process audit first, tools second. The reverse order ends with buying a solution and hunting for a problem to fit it.
The procedure: five steps you can run on paper
The whole exercise comes down to breaking work into tasks and calculating how much of each can realistically be lifted off a person. Work through the steps in order, and do not skip the third one, because that is what separates an honest calculation from a wishful one.
- List the repeatable tasks of a month, not departments. A task has a start, an end, and a checkable output: "listing a new delivery in channel X", not "marketplace operations". Sources: calendars, the helpdesk queue, the spreadsheets someone opens every week, and half an hour with each person who actually does the work.
- Assign each task a realistic monthly time. Calculate it as frequency times the length of one run, plus rework and context switching. If you can, measure for two weeks instead of estimating from memory. Memory systematically understates routine and overstates firefighting, because firefighting is what gets remembered.
- Estimate the rule-based share. Ask: how much of this task could a new hire complete given only a written instruction and the data, with no questions about context? That is the rule-based share, and only that share is a candidate for automation. The rest is judgment, and judgment does not evaporate.
- Subtract residual oversight. Someone will review outputs, handle exceptions, and react to changes in the underlying systems. Early on that oversight is substantial, and it shrinks only once you have evidence the workflow is stable. Planning for zero oversight is the most common error in these calculations.
- Only the difference is a saving. Record it as a dated assumption with the figures you used, not as a fact. In a quarter you will see how far off you were, and that is a legitimate output of the audit.
Step three is the hard one, because it demands honesty about your own process. The natural instinct is to call the whole task rule-based, since "we do it the same way every time". Usually we do not: in a product description, the structure and the attribute fill are rule-based, but the decision about what to emphasize on a flagship product is not. We showed what that split looks like across concrete tasks, and how much stays with the human in each, in a separate piece on the anatomy of automation savings. Here the principle is enough: you automate part of a task, not a task.
The second criterion: error cost and sales impact
Hours alone will not set your sequence. The task with the biggest hour count can be the worst first choice if its errors are expensive or hard to detect. So give every candidate two more scores, each on a scale of 1 to 3. Error cost: does a bad output stay inside your team, or does a customer or a marketplace see it, and can it be reversed in a minute or does it require corrections in several systems? Sales impact: does the task sit on the path to revenue, for instance by holding up the publication of new stock, limiting channel coverage, or slowing responses to pre-sales questions?
These two scores do not decide whether a task is suitable for automation at all, because the qualification rule from the previous chapter already settled that. They decide the order. A task with high sales impact and low error cost goes to the front of the queue even if it carries fewer hours than others. A task with high error cost goes further back, not because it is impossible, but because it needs more oversight and better data, and both are easier to build once you already have one working implementation behind you.
Three questions that set priority faster than any scoring sheet: does the task return at least weekly, can a bad output be reversed in a minute, and can you name the single number that should move after implementation? Three yeses make a first-implementation candidate. One no puts it in the queue. Two nos mean the task stays with a person for now.
The implementation queue: what goes first, what goes second
The output of the audit is not a ranking but a queue with three tiers. First comes a task with high repetition, a large rule-based share, cheap errors, and an obvious measure of success. You will usually find it in the catalog or in feeds, because there you work on data you already hold. Second is the higher-impact task that needs integrations, system access, or a data cleanup first. You start preparing it in parallel but launch it after the first, using what the first one taught you. Third is the waiting room: high impact combined with high error cost, typically anything that touches direct communication with customers.
One rule organizes this queue better than all the scores combined: do not launch three implementations at once. Not because of budget, but because of attribution. When three things start in the same month and operational results do not improve, you cannot tell which one failed, and the whole company concludes that "AI does not work here". With one implementation at a time, every outcome, good or bad, has a clear cause. The audit also produces something else worth having: a list of tasks you have consciously decided not to automate. That is a real result, because it closes the question and saves you the next round of debate.
The baseline: measure before you start
The most common reason an AI implementation ends without a verdict is mundane: nobody wrote down the state before it started. A quarter later you can no longer reconstruct how long the work used to take, so the discussion about impact collapses into impressions. So record four things for the task going first: monthly hours from step two, a quality measure (reworks, channel rejections, complaints in that category), a speed measure (time from the moment the work becomes possible to the moment it is done), and who does it today.
Add the context, because without it a quarterly comparison means nothing: number of items, orders, channels, and languages, plus the date of measurement. If the catalog grows by half in the meantime, a drop in hours means something different from what it appears to mean. Take the follow-up measurement thirty days after launch, not after a week: the first days always look worse, because tuning is still in progress and oversight is at its highest. If you want to place the calculation inside total cost rather than hours alone, run it through the free Blueprint Check calculator, which computes automation savings from your own assumptions.
Keep the first implementation small
The scope of the first implementation is a strategic decision, not a question of ambition. The recommendation is unambiguous: one process, one owner on your side, one number that should move, and a horizon measured in weeks rather than quarters. Not "AI transformation of the company", but "cut the time from new delivery arriving to being live in two channels". The reason is practical: a small scope quickly reveals two things no presentation can predict, namely whether your data is good enough and whether the team can carry the oversight inside its normal working day.
A large scope takes both of those away from you. When an implementation spans five processes at once, failure cannot be attributed and success cannot be verified, so the company ends up with an opinion instead of a conclusion. A small start has one more side effect, usually the most valuable: after the first implementation the team can describe in its own words what AI does well and what it is not allowed to touch, and the decisions that follow stop being a matter of belief. The audit almost always ends in the same discovery: the constraint is neither the models nor the budget, but the state of the data they would work on. That is the subject of the next chapter.
Questions
How long does an operations audit take before an AI implementation?
In a small team the task list comes together in one afternoon, and assigning realistic times usually takes about two weeks if you measure rather than estimate from memory. Treat that as orientation, not a norm: with several people and many channels it takes longer. What matters more than speed is that the timings come from the people who actually do the work.
Who should run the process audit?
The people doing the work supply the timings and the exceptions, and one decision-maker owns the priorities and the number that should move. An audit run purely at management level systematically understates routine work, because routine is invisible from above. An audit run purely from the bottom up usually fails to set the sequence, because it lacks the commercial weighting.
What if I have no data on how long tasks take?
Measure for two weeks instead of guessing. A simple log is enough: task, date, start and end time. If measuring is impossible, calculate frequency times the length of one run, add an allowance for rework, and label the result explicitly as a dated assumption. A labelled assumption is useful because it can be corrected later. A number quoted without a source stays in the deck forever.