By the end of this lesson you will be able to
- Establish a baseline before you change anything
- Pick metrics that connect to what the business already cares about
- Report results in a form a sceptical CFO would accept
The most common measurement mistake is tracking activity instead of outcomes. "Sixty percent of staff used AI this week" tells you nothing about whether anything got better. It is the easiest number to produce and the least persuasive one to present.
Baseline first, or do not bother
You cannot measure improvement without a starting point, and reconstructing one afterwards from memory produces numbers nobody believes — including you. Measure for at least four weeks before anything changes. Shorter than that and normal week-to-week variation swamps the signal.
For each process you intend to improve, record:
Time on the target task
Sample from several people. Do not rely on a single estimate, and do not rely on estimates at all where you can log the real thing — people are consistently wrong about how long routine work takes, usually underestimating.
Current output quality
Use whatever measure you already have: error rate, revision rate, first-draft-to-approval ratio, client satisfaction. Inventing a new quality metric for this is a warning sign that you will not keep measuring it.
Volume
Proposals per month, tickets per day, reports per week. Volume is what turns a time saving into a capacity story, which is the version leadership responds to.
Cost per unit, where it applies
What does one of these currently cost to produce? Not every process has a sensible answer; skip it where it does not.
Four tiers of metric, in ascending order of credibility
| Tier | Example | Credibility | Catch |
|---|---|---|---|
| Activity | Weekly active users | Lowest | Proves nothing improved |
| Time | Hours per week on the target task | Moderate | Only counts if the time is reinvested |
| Volume / quality | Proposals per month, revision rate | Good | Needs a real baseline |
| Business outcome | Proposal win rate, time to hire, rework rate | Highest | Hardest to attribute cleanly |
Activity metrics are worth watching internally as an early warning that a rollout is dying. They are not worth presenting as results, and presenting them tends to damage your credibility on the ones that matter.
The time-savings trap
"We saved four hours a week" is the most-quoted AI metric and the least meaningful on its own. Saved time only becomes value if it goes somewhere. If a bookkeeper saves three hours and spends them on the same backlog they always had, nothing changed except how the backlog was processed.
A number that gets challenged
- "AI saves us 40 hours a month"
- No baseline, estimated after the fact
- No evidence the hours went anywhere
- Quality impact unmeasured
A number that holds up
- "Quote turnaround went from 3 days to same-day"
- Baselined over four weeks beforehand
- Volume up 20% with the same headcount
- Revision rate down from 1 in 4 to 1 in 9
Count the whole cost, not the subscription
An ROI calculation that only counts licence fees will overstate the return, and someone will eventually notice. Include:
- Subscriptions and API spend — usually the smallest line.
- Setup and integration time, at a real loaded hourly rate.
- Ongoing verification. This never goes to zero for work that matters, and it is the line most often left out.
- Training and the productivity dip while people learn.
- Rework when a model update quietly changes behaviour.
What to do when the answer is no
Sometimes the measurement says it did not work. That is a useful result and should be reported as plainly as a positive one — a team that only ever reports wins is a team nobody believes on the third project.
Before concluding the technology failed, check the three things that usually failed instead:
- Was the process actually defined? An undefined process automated produces fast inconsistency, not improvement.
- Did people know what to use it for in their role? Low adoption is a training result, not a technology result.
- Was the target task a good fit at all? Recall-heavy work was never going to work well. That is a selection error.
How long before we should expect a measurable result?
Sixty to ninety days for anything behavioural. Time savings on a single well-defined task can show inside four weeks. Business outcome metrics like win rate need a full sales cycle, so be honest up front about when you can answer, or you will be asked at thirty days and have nothing.
How do we attribute a business outcome to AI specifically?
Cleanly, you often cannot, and claiming otherwise invites an argument you will lose. Say what you can defend: the process metric moved, here is the baseline, here is what else changed in the period. Directional honesty survives scrutiny better than a precise number that does not.
Is a control group worth it?
For a team of eight, no — the overhead exceeds the value. For a rollout across several similar teams, staggering the launch gives you a natural comparison almost for free, and it makes the result substantially more credible.
Key takeaways
- Baseline for four weeks before you change anything, or the result will be arguable forever.
- Track outcomes, not activity. Usage rates are an early warning, not evidence.
- Saved time only counts if it demonstrably went somewhere.
- Count verification and integration in the cost, not just the subscription.
- Report at 30, 60, and 90 days — and report the negative results too.
