Resources / Research & numbers

How to prove an AI project worked: measure before you build

The short answer: record the number before anything is built. One workflow, one metric, measured the way the work actually runs, written down and signed by both sides. Do that and “did it work?” has an answer on day 90. Skip it and the question becomes a matter of opinion — and opinions lose budget arguments.

That is the whole of it. The rest of this explains why it is so often skipped, what the research actually says, and how to record a baseline in two weeks without spending anything.

What the research measures — and what it doesn’t

Three numbers get quoted constantly. They are worth reading carefully, because they do not say what people assume.

FindingSource
95% of enterprise GenAI pilots produced no measurable P&L returnMIT Project NANDA, 2025
42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year beforeS&P Global
76% of small businesses use AI; 14% have it fully embeddedGoldman Sachs, 2026

Read quickly, that looks like evidence the technology does not work. Read again. “No measurable return” is not the same as “no return.” It is a statement about measurement.

The Goldman figure is the clearest. The gap between 76% using AI and 14% having it embedded is not a gap in capability — the same models are available to everyone. It is the gap between trying something and changing how the work runs.

How a project ends up unprovable

The sequence is almost always the same, and nobody in it does anything wrong.

Someone senior sees a demo and gets interested. A pilot is scoped, usually around the workflow that is easiest to automate rather than the one that costs the most. Something gets built. It half-works, which is normal for a first version. The person who championed it moves on to the next quarter’s problem.

Six months later a director asks whether it saved anything, and the room goes quiet. Not because the answer is no — often nobody knows — but because there is nothing to compare against. The subscription keeps renewing. At the next budget round it is cut, and it goes into the company’s memory as “we tried AI and it didn’t work.”

The failure was not in the build. It was on day one, when nobody wrote down what things were like beforehand.

Why nobody records the baseline

Four reasons, in rough order of how often they come up.

It feels like delay. Two weeks of measuring produces nothing you can demonstrate. The pressure is to show progress, and a working prototype looks like progress in a way that a spreadsheet of current cycle times does not.

Everyone believes they already know the number. Ask a sales director how quickly enquiries get answered and you will get an answer immediately. Measure it and the answer is usually wrong, because the average hides the ones that sat over a weekend.

The data is scattered. The number lives across an inbox, a CRM and someone’s memory. Assembling it is genuinely tedious.

Nobody wants the number on record. This is the real one, and it is rarely said out loud. A written baseline is also a written account of how bad things currently are, and somebody owns that process.

That last reason is why the baseline has to be signed by both sides. It stops being an accusation and becomes the starting line.

How to record one in two weeks

1. Pick one workflow, not a category

“Customer service” is not a workflow. “Time from a customer email arriving to a human replying” is. If you cannot describe it as something that starts, happens and finishes, it is too big to measure.

2. Pick one metric, not a dashboard

One number, chosen because moving it is worth money. Cycle time, cost per task, error rate, throughput per person. Twelve metrics on a dashboard means nobody is accountable for any of them, and it guarantees that at day 90 you can find something that improved.

3. Measure the way the work actually runs

Not the way the process document says it runs. Sit with the people doing it for an afternoon. The gap between the documented process and the real one is usually where the money is going.

4. Take the median and the worst 10%, never the average

This matters more than any other line here. Averages hide the failures that cost you customers. If most enquiries get a reply in two hours but one in ten sits for three days, the average looks respectable and the business is bleeding from the tail.

5. Measure for two weeks minimum

One week catches a fluke. Two weeks catches the pattern, including the Monday pile-up and the Friday drop-off.

6. Write it down and both sides sign it

Dated. Naming the metric, the method, and who took the measurement. This is the document that makes the day-90 number checkable rather than arguable.

What a baseline document contains

Short, and boring on purpose:

Five items. It fits on two pages. It is the least interesting thing an engagement produces and the only thing that makes the rest provable.

The arithmetic it lets you do

Once you have a baseline you can answer the question that actually decides budget: how much would this number have to move to be worth doing?

If a workflow takes 40 hours a week across a team, and a change would remove a quarter of that, you are looking at 10 hours a week. Put your own loaded cost per hour against it and you have a payback period rather than a hope. If the arithmetic does not work, you have saved yourself the build — which is a good outcome, arrived at cheaply.

You cannot do any of that arithmetic without the first number.

The uncomfortable part

Recording a baseline means agreeing in advance what failure would look like. Most projects avoid this, which is precisely why most projects cannot prove they succeeded. If you are only willing to measure things that will flatter you, you are not measuring — you are marketing.

Two weeks and no money is what it costs. The software costs real money. Doing them in that order is most of the trick.


This is the first phase of every engagement we run — see how it works for what the other two produce, and what gets published for the format the day-90 result is handed over in. The AI Impact Diagnostic is this, done in one to two weeks, and its fee comes off the build if you go ahead.