AI ROI is almost always presented the same way: one slide, one big number, a savings percentage and an arrow pointing up. And it is almost always wrong. Not because the person presenting it is lying, but because half the equation is missing: there is no baseline to compare against, the cost only covers the licence, and the savings are counted in hours nobody has actually stopped paying for. That number does not survive the first uncomfortable question from the board, and the project's credibility goes down with it.
This guide is the method we use to calculate the return on an AI project so that it holds up to an internal audit: what to measure before you start, which costs genuinely count, how to tell hours freed up from euros saved, and what to do with the benefits that will never fit in a spreadsheet. No magic formulas, no percentages borrowed from third-party reports.
What is AI ROI and why is it usually miscalculated?
The formula is no mystery: net benefit divided by total investment. The problem is never the division, it is the two numbers feeding it. In AI projects they go wrong for five reasons that repeat with suspicious regularity:
- There is no baseline. Nobody measured how long it took before, at what error rate and at what cost. Without that starting point, any improvement is a well-written opinion.
- Cost means the licence only. Integration, data cleanup, internal team time and second-year maintenance all quietly disappear.
- Savings are counted in theoretical hours. "We save 20 hours a month" is not money if those 20 hours are still on the payroll doing something else nobody has decided on.
- Everything that improved gets attributed to the project. If a new salesperson joined the same quarter and conversion went up, it was not all the model.
- It is measured too early. The first six weeks of any tool are a bad moment to judge it: the learning curve penalises and the novelty effect exaggerates.
An ROI you cannot reconstruct six months later from data that already existed is not a measurement, it is a narrative.
The practical consequence of getting it wrong is not just an optimistic slide. It is that the second project gets approved on a false basis, and when the third one underdelivers, management concludes that "AI doesn't work here" when what did not work was the measurement method.
What do you need to measure before you start?
The baseline is the cheapest part of the project and the one most people skip, because it demands two weeks of boring work before the fun begins. It is non-negotiable: if you do not have it before you switch anything on, you will not be able to reconstruct it afterwards.
For any candidate process, measure four things over two to four representative weeks:
- Volume: how many units get processed (invoices, tickets, orders, queries) per week, and their seasonality.
- Time per unit: timed on a real sample, not estimated in a meeting. Estimates from memory are systematically wrong, in both directions.
- Error and rework rate: how many units come back, what each round trip costs and who pays for it.
- Cost per unit: the loaded hourly cost of whoever does the work multiplied by actual time, plus the cost of errors.
Add a fifth if the process touches the customer: response time. In many cases the real return is not in the savings, it is in quoting in two hours instead of two days — and you can only prove that if you have the "before" figure.
One detail that saves arguments later: keep the sample. Not the summary, the sample. When someone challenges the calculation a year from now, being able to open the 200 invoices you measured in March and see the recorded times is worth more than any presentation.
Which costs genuinely belong in the calculation?
The three-year total cost of ownership of an AI project rarely looks like the initial budget, and the gap is not bad faith: it is that only the invoiced parts get budgeted. This is the list we use, and it is worth filling in completely even if some rows come out at zero.
| Line item | What it covers | Where it gets underestimated | |---|---|---| | Licences and usage | Subscriptions, model usage costs, infrastructure | Usage grows with adoption; year 1 does not predict year 2 | | Implementation | Analysis, development, ERP/CRM integration | Integrations eat more budget than the model does | | Data preparation | Cleanup, history, fixing master records | The line that overruns most and gets budgeted least | | Internal time | Business hours spent defining, testing and validating | Not invoiced, but it is time not spent on something else | | Training and adoption | Sessions, documentation, hands-on support | Without it the tool gets half-used and the return never arrives | | Maintenance | Tuning, retraining, changes in source systems | Budget 15 % to 25 % of implementation per year | | Human oversight | Reviewing edge cases, quality control | In a critical process this line never disappears |
Two notes on this table. First: if a vendor gives you a price you cannot break down into these lines, you are not comparing offers, you are comparing headline figures — we cover that in what an AI project really costs. Second: human oversight is the most forgotten cost and the one that decides whether a use case makes sense at all. A model that is 92 % accurate but requires reviewing 100 % of its output saves almost nothing.
Hours freed up or money actually saved?
This is the most common trap in the calculation, and it is worth being blunt: an hour freed up is only money saved if that hour stops costing something or turns into revenue. Everything else is capacity, which is valuable but not the same thing.
Three scenarios and what you can honestly book in each:
| Situation | Does cash change? | What you can book | |---|---|---| | External or temporary hours are no longer contracted | Yes, directly | Real saving, with an invoice that proves it | | Growth is absorbed without adding headcount | Yes, avoided | Cost avoidance, if the growth is real and hiring was planned | | The team spends those hours on higher-value work | Not immediately | Reassigned capacity; justified by the results of that work | | Costly errors are reduced | Yes | Real saving, measurable against the rework baseline | | People work under less pressure | No | Qualitative benefit; report it separately, do not monetise it |
The rule we apply with management is simple: only monetise what you can point to in a P&L or in an invoice that no longer arrives. Everything else goes in a separate section of the report, described honestly. A defensible 60 % ROI is worth infinitely more than a 300 % that falls apart when the CFO asks where that money actually is.
There is one legitimate exception worth naming: incremental revenue. If a prioritisation model lets sales reach more opportunities with the same headcount and you can attribute the difference using a control group, that is euros. With a control group — not with "we've been selling more since we started using it".
Over what period do you measure, and against what?
A reasonable horizon for a contained use case is six to twelve months from the moment it is in production, not from the moment the contract is signed. And the honest comparison is rarely "before and after": it is a design that isolates the project's effect from everything else that happened in the company that quarter.
In order of reliability:
1. Control group: one team, branch or segment keeps working as before for eight to twelve weeks. It is the cleanest option and almost always feasible if you plan for it up front. 2. Staged rollout: it goes live in phases and you compare cohorts that do not have it yet. Useful when you cannot keep a group without the tool for long. 3. Before and after with adjustment: compare against the same period last year and explicitly discount what changed (volume, headcount, prices). It is the weakest option, but saying so openly already makes it more rigorous than most reports.
Set the review calendar before you start too: a three-month check (usage and quality indicators), a six-month one (first economic calculation) and a twelve-month one (consolidated figure with real maintenance costs). This connects with the discipline of choosing KPIs people actually use: every review has an owner, a date and a number compared against the baseline.
One warning about timing: the first three months usually show a worse return than the project will eventually deliver. The adoption curve weighs heavily, and if you kill the project in week eight over a weak number you will be killing cases that work. It is the pattern we describe in why AI pilots fail.
What do you do with benefits you cannot put in euros?
They exist, they are real, and pretending they can be monetised helps nobody. The serious way to handle them is to give each one its own indicator, target and review date, even if it never enters the ROI division:
- Risk reduction: fewer errors in a regulated process, traceability of who approved what. Measured in incidents avoided and in response time when an audit request lands.
- Service quality: response time, percentage of queries resolved first time, satisfaction measured the same way before and after.
- Knowledge retention: procedures that stop living in one person's head. You notice it when they go on holiday.
- Room to grow: absorbing 30 % more volume without redesigning the process. It gets proven the day the peak arrives, not before.
- Organisational learning: the first project leaves behind organised data, written criteria and trained people who make the second one cheaper. It is the honest argument for approving a case with a modest return, and it ties directly to training your team.
Put them on a second page of the report, each with its metric and review date. Management appreciates the distinction more than people expect: separating cash from capacity is exactly the signal that the rest of the numbers were done properly.
Frequently asked questions
What is a reasonable ROI for an AI project?
It depends on the type of case, but for contained automations with high volume it is common to recover the investment in six to eighteen months. Be sceptical of any promise of a return in weeks, and equally of projects that cannot say which month they expect to break even: if the vendor will not commit to a range, they have not done the calculation.
Can you measure the ROI of a general-purpose AI assistant?
It is the hardest case, because usage is spread across many people and many tasks. The practical way out is to pick two or three concrete, measurable uses — drafting a specific type of document, answering a specific type of query — and measure only those against a baseline, rather than trying to capture the diffuse benefit across the whole organisation.
When should you stop a project for insufficient ROI?
When the adoption period is over, usage indicators are low and the cost per unit has not improved against the baseline. Usage is the early signal: if people are not using it there is no possible return, and the problem is usually process or fit, not the model. Stopping in time and documenting why is a healthy decision, not a failure.
Who should calculate the ROI, the vendor or the company?
The company defines the calculation and the vendor feeds it with verifiable data. If whoever has to prove the value also picks the metrics and the comparison period, the outcome is predictable. A good partner proposes the method before starting and accepts that the client sets the baseline.
If you have an AI project running and you would struggle to defend its return in front of the board, the problem is rarely the calculator: it is a missing baseline and too much noise in the cost side. In the audit we review which processes you actually measure, which data lets you reconstruct a before and after, and which use cases have a genuinely defensible return. If you would rather start with a conversation, let's talk for half an hour and we will tell you straight whether your case adds up yet.
Shall we apply it to your case?
The 360° AI Audit turns these ideas into a concrete plan for your company: three weeks, fixed price and the full picture of your AI before spending a euro.
See the 360° Audit→ Let's talk↗