Measuring AI ROI without lying to yourself
Most AI business cases are built on hours saved that never left the payroll. A method for measuring what actually changed — baselines, counterfactuals, and the four categories of real return.
Most AI business cases are arithmetic performed on a fiction.
The fiction goes: the agent handles 8,000 tickets a month, each ticket used to take 6 minutes, that's 800 hours, at €25 fully loaded that's €20,000 a month, €240,000 a year, enormous ROI, ship it.
Then twelve months pass and the support team is the same size, the budget is unchanged, and nobody can find the €240,000.
The arithmetic isn't wrong. The premise is. Here's how to build a business case that survives the second year.
The core distinction: hours saved vs. money moved
Hours saved are real. They're just not automatically money.
An hour freed becomes financial return in exactly three ways:
- Headcount you don't hire. Growth that would have required a new hire, absorbed instead. Real, and the most common genuine return.
- Headcount you reduce. Real, immediate, and the one most operators don't want and shouldn't pretend to.
- The hour is redeployed to something that generates value. Real, but only if you can name what that something is and measure it.
If none of the three applies, the hour was saved and nothing happened. The team is less stressed — which is a real and worthwhile outcome — but it isn't a line in the P&L, and presenting it as one is how AI programmes lose credibility in year two.
Be explicit about which mechanism applies, per workflow, before you build. "We will not hire the two seasonal support agents we hired last Q4" is a claim you can check in January. "800 hours saved" is not.
Establish the baseline before you build
This is the step that gets skipped, and skipping it makes everything afterwards unfalsifiable.
Before a single line of code, record:
- Volume — units of work per period, with seasonality
- Time per unit — measured, not estimated. Estimates run 30–50% optimistic in our experience.
- Fully loaded cost per hour — salary, employer costs, tooling, management overhead, recruitment amortised. Usually 1.4–1.7× base salary.
- Quality baseline — error rate, CSAT, rework rate, SLA compliance
- The current failure mode — what goes wrong today, how often, what it costs when it does
That last one is the most valuable and the most neglected. A large share of real AI return isn't hours at all — it's errors that stop happening. A wrong-price ad that ran for six days. A stockout nobody caught. A checkout bug detected in four minutes instead of four hours. Those have euro values, and they're usually larger than the labour line.
The four categories of real return
1. Cost avoided. The clearest. Growth absorbed without hiring, agency retainers not renewed, contractor hours not booked, overtime not paid. Test: would we genuinely have spent this? If you can point to a specific budget line or a hiring plan that changed, it's real.
2. Revenue enabled. Recovered carts, saved cancellations, faster speed-to-lead, more SKUs launched per quarter, more markets served. Higher variance and harder to attribute — you need a counterfactual (see below) — but this is usually the larger number, and the one that gets a programme funded rather than tolerated.
3. Loss prevented. Errors caught before they cost money. Ad spend on a broken creative. A margin-negative discount that would have run all weekend. Compliance exposure. Chargebacks won that would have been lost. Quantify the historical rate and the average cost per incident, then measure the new rate.
4. Speed. Where cycle time has a value. Getting to market four weeks earlier is worth four weeks of that product's contribution. Answering a lead in 60 seconds rather than 6 hours has a measurable conversion differential you can look up in your own CRM.
Categories 3 and 4 are consistently under-counted, because nobody has a line item for "the disaster that didn't happen."
Counterfactuals, or the attribution problem
You launched a cart recovery agent in March. April revenue is up 14%. How much was the agent?
Without a counterfactual, unknowable. April also had a campaign, better weather, and an easier comparison base.
Three ways to get one, in descending order of rigour:
Holdout groups. Withhold the treatment from a random 10% of eligible traffic. Compare. This is the gold standard and it costs you 10% of the benefit — which is nearly always worth paying for a number you can defend in a board meeting.
Staged rollout. Deploy market by market or segment by segment. Untreated segments are your control. Weaker than randomisation because segments differ systematically, but often the only politically feasible option.
Pre/post with a control metric. Compare the period before and after, alongside a metric that should not have been affected. If recovered revenue rose 40% and overall site conversion was flat, the recovery agent is the plausible cause. Weakest of the three; still far better than nothing.
Whatever you choose, decide it before launch. A counterfactual designed after you've seen the results is a rationalisation.
Count the full cost
Business cases that only count model spend are understating by a wide margin.
Build: engineering time, integration, data preparation (usually the largest and most underestimated line), review cycles.
Run: model and infrastructure spend, monitoring, the human hours still involved — escalation handling, spot-checking, approval queues.
Maintain: the one everyone forgets. Agents drift when the business changes. New products, new policies, new edge cases, model deprecations, provider API changes. Budget 15–25% of build cost annually as ongoing maintenance, or you'll discover it as an unplanned project in month nine.
Change: training, process redesign, documentation, the productivity dip while people adjust. Real, temporary, and reliably absent from business cases.
A worked example
A D2C brand automating support. Honest version:
Baseline: 25,000 tickets/month. Median handle time 5.2 minutes (measured from timestamps, not asked). Fully loaded €26/hour. 11 FTE. Growth trajectory implies 3 more hires within 12 months.
After 6 months: 71% automated. Median handle time on the remaining 29% rose to 9 minutes — the easy tickets left, so the residue is harder. Escalation precision 94%.
The return:
- Cost avoided: 3 hires not made. €124,000/year. Real — the hiring plan is documented and was cancelled.
- Cost avoided: seasonal contractor spend, €31,000/year. Real.
- Revenue enabled: two support leads moved to retention; the flows they built are measured against a holdout at €340,000 incremental/year. Real, with a defensible counterfactual.
- Loss prevented: faster detection of a checkout error, historically ~2 incidents/year at ~€40,000 each. Partly real — estimate based on 3 years of incident history, flagged as an estimate.
- Hours saved on remaining team: ~600/month. Not counted. Nobody left, nothing was redeployed. Reported as a wellbeing outcome, not a financial one.
The cost: €78,000 build, €31,000/year run (model spend, infra, escalation hours), €16,000/year maintenance.
Year one: €495,000 real return against €125,000 total cost. Roughly 4× — a fraction of what the naïve calculation would have claimed, and a number that survived the board asking where it came from.
The four questions to ask any business case
- Which specific budget line changes, and when? If nobody can name one, it's an efficiency story, not a financial one. That's fine — say so.
- What's the counterfactual, and was it designed before launch?
- Does the cost include maintenance and the human hours still involved?
- What happens in month nine when the business changes?
A business case that answers all four honestly will show a smaller number than the fantasy version. It will also still be true in twelve months — which is the only property that matters, because the second AI project gets funded on the credibility of the first.
The uncomfortable one
Sometimes the honest answer is that a workflow isn't worth automating. Volume is too low, variance is too high, the process should be deleted rather than accelerated.
We tell clients this during the audit, before there's a contract to protect, and it's the reason we do audits as a separate paid engagement. Automating a process that should have been eliminated is the most expensive way to make a process faster.
Our audit produces a ranked automation map with hours and euros attached — yours to keep whether or not you hire us. Book one.