Blog

How to tell whether your automation is actually working

An agent or automation that nobody measures is one you are trusting on feel. Here is what to write down before you switch it on, what to count once it is running, and how to spot it going quietly wrong.

The most common thing we see after a business switches on an automation is nothing. Not failure: nothing. It runs, nobody complains, and six months later nobody can say whether it saved an hour a week or cost them three customers. "It seems fine" is not a measurement, and an agent that speaks to your customers deserves better than that. Here is how we keep score, in the order you need to do it.

Write down what "working" means before you switch it on

Measurement starts before the thing exists. Before an agent or automation goes live, write one sentence about what it is for and one number you expect to move. "Follow up every quote that has had no reply for five days, so fewer quotes go quiet" is a purpose. "Right now about a third of quotes never get a reply" is the number you are trying to shift.

Then take the baseline. Count the thing as it is now, for a few weeks if you can: missed calls, no-shows, hours spent on the task, days to get paid. This is the step people skip, and skipping it means that whatever happens later, you will not be able to say what changed. If you have no baseline, the honest answer to "is it working?" is always "we don't know".

Count outcomes, not activity

Automations love to report how busy they are. Messages sent, emails read, records updated. None of that is the point. The point is what the business got: quotes that turned into jobs, appointments that were kept, invoices that got paid sooner, evenings that were not spent re-typing.

So for each automation, pick one outcome number and one or two activity numbers that sit behind it. For a quote chaser, the outcome is the share of quotes that get a reply; the activity is how many follow-ups went out and how many customers responded to them. Activity tells you it is running. Outcome tells you it is worth running.

Track the misses as carefully as the hits

An agent that does the right thing most of the time is only useful if you know what happens the rest of the time. Four things are worth counting every week:

  • Hand-offs: how often it gave up and passed the job to a person, and why.
  • Corrections: how often a person changed what it had drafted before it went out.
  • Overrides: how often a person stopped it doing something it was about to do.
  • Complaints: anything a customer said that suggests they noticed it was a machine, or noticed it being wrong.

The first three are easy to log if the agent is built with a person in the loop, which is how we recommend starting anyway. A rising correction rate is the earliest warning you get that something has changed and the brief no longer fits.

Watch the cost per job, not the monthly bill

An automation has a running cost: software subscriptions, model usage if there is an AI in it, and the time people spend checking it. The useful way to look at that is per job done, not per month. Divide what it costs by what it did, and compare that with what the same job cost when a person did it.

Model usage in particular can creep. An agent that starts taking more steps to do the same task, or that is being asked to handle more than it was designed for, will show up in the bill before it shows up anywhere else. A monthly glance at cost per task catches that early.

Read the log

Numbers tell you that something is off. The log tells you what. Every agent we build keeps a record of each run: what it saw, what it decided, which tools it used, what it sent. Once a week, someone reads a handful of them, picked at random rather than the ones that were flagged.

This takes twenty minutes and it is the single most useful habit on this list. It catches the reply that was technically correct and slightly rude, the customer it kept calling by the wrong name, the case it handled perfectly that nobody had thought to design for. You cannot find those in a dashboard.

Expect it to drift

Nothing about your business stays still. Prices change, a service gets dropped, the person who used to approve things leaves, a supplier renames a product. The agent does not know any of that unless it is told, and it will carry on confidently using last year's facts.

Put a date in the diary to reread the agent's instructions and the information it draws on, every quarter at least, and every time something about the business changes. Treat it like the price list on the website: if you would update one, update the other.

A monthly review that fits on one page

Pull all of this into a short review once a month. It needs five lines per automation:

  1. The outcome number, this month against the baseline.
  2. The activity numbers, so you can see it is running as expected.
  3. Hand-offs, corrections and overrides, and whether they are going up or down.
  4. Cost per job.
  5. One thing you noticed reading the log.

If the outcome has moved in the right direction and the misses are steady or falling, leave it alone. If the outcome has not moved, the automation is either not the right one or is not being used, and either is worth knowing. If the misses are climbing, the brief needs updating before anything else.

The rule underneath all of this

Do not automate anything you are not prepared to measure. That sounds strict, and it is meant to be. The whole case for handing a job to software is that it does the job better or cheaper than before. If you cannot show that, you have not automated the job, you have just stopped looking at it.

If you have an automation running and no idea whether it is earning its keep, book a chat and we will help you work out what to count.

More from the bench.

The automations most small businesses start with

The handful of jobs that small firms most often hand over to software first, what each one actually involves, and how to pick the one that will pay for itself soonest.

What an AI agent is actually made of

Strip away the marketing and an AI agent is six ordinary parts working together. Here is what each one does, why it matters, and which one usually goes wrong.

Free · No obligation

Got a question this didn't answer?

Thirty minutes, free, no obligation. Tell us the three things eating your week and we'll tell you which one we'd automate first.