Rolling AI Out Lesson 2 of 3

Measuring Whether It Actually Worked

Usage statistics are not evidence. Here is what to baseline, what to track, and what to say to a sceptic.

Applied Evergreen 8 min

Worth reading first: Rolling AI Out to Your Team

By the end of this lesson you will be able to

  • Establish a baseline before you change anything
  • Pick metrics that connect to what the business already cares about
  • Report results in a form a sceptical CFO would accept

The most common measurement mistake is tracking activity instead of outcomes. "Sixty percent of staff used AI this week" tells you nothing about whether anything got better. It is the easiest number to produce and the least persuasive one to present.

Baseline first, or do not bother

You cannot measure improvement without a starting point, and reconstructing one afterwards from memory produces numbers nobody believes — including you. Measure for at least four weeks before anything changes. Shorter than that and normal week-to-week variation swamps the signal.

For each process you intend to improve, record:

  1. Time on the target task

    Sample from several people. Do not rely on a single estimate, and do not rely on estimates at all where you can log the real thing — people are consistently wrong about how long routine work takes, usually underestimating.

  2. Current output quality

    Use whatever measure you already have: error rate, revision rate, first-draft-to-approval ratio, client satisfaction. Inventing a new quality metric for this is a warning sign that you will not keep measuring it.

  3. Volume

    Proposals per month, tickets per day, reports per week. Volume is what turns a time saving into a capacity story, which is the version leadership responds to.

  4. Cost per unit, where it applies

    What does one of these currently cost to produce? Not every process has a sensible answer; skip it where it does not.

The trapStarting the pilot and deciding to measure afterwards. By then the honest baseline is gone and every number you produce is arguable. Four weeks of boring measurement up front is what makes the eventual result defensible.

Four tiers of metric, in ascending order of credibility

TierExampleCredibilityCatch
ActivityWeekly active usersLowestProves nothing improved
TimeHours per week on the target taskModerateOnly counts if the time is reinvested
Volume / qualityProposals per month, revision rateGoodNeeds a real baseline
Business outcomeProposal win rate, time to hire, rework rateHighestHardest to attribute cleanly
Pick from as far down this table as you can honestly reach.

Activity metrics are worth watching internally as an early warning that a rollout is dying. They are not worth presenting as results, and presenting them tends to damage your credibility on the ones that matter.

The time-savings trap

"We saved four hours a week" is the most-quoted AI metric and the least meaningful on its own. Saved time only becomes value if it goes somewhere. If a bookkeeper saves three hours and spends them on the same backlog they always had, nothing changed except how the backlog was processed.

A number that gets challenged

  • "AI saves us 40 hours a month"
  • No baseline, estimated after the fact
  • No evidence the hours went anywhere
  • Quality impact unmeasured

A number that holds up

  • "Quote turnaround went from 3 days to same-day"
  • Baselined over four weeks beforehand
  • Volume up 20% with the same headcount
  • Revision rate down from 1 in 4 to 1 in 9

Count the whole cost, not the subscription

An ROI calculation that only counts licence fees will overstate the return, and someone will eventually notice. Include:

  • Subscriptions and API spend — usually the smallest line.
  • Setup and integration time, at a real loaded hourly rate.
  • Ongoing verification. This never goes to zero for work that matters, and it is the line most often left out.
  • Training and the productivity dip while people learn.
  • Rework when a model update quietly changes behaviour.
Reporting cadenceReport at 30, 60, and 90 days. Thirty is too early for outcomes but the right time to catch a rollout dying. Ninety is the first honest read on whether behaviour actually changed rather than people trying something new.

What to do when the answer is no

Sometimes the measurement says it did not work. That is a useful result and should be reported as plainly as a positive one — a team that only ever reports wins is a team nobody believes on the third project.

Before concluding the technology failed, check the three things that usually failed instead:

  • Was the process actually defined? An undefined process automated produces fast inconsistency, not improvement.
  • Did people know what to use it for in their role? Low adoption is a training result, not a technology result.
  • Was the target task a good fit at all? Recall-heavy work was never going to work well. That is a selection error.
How long before we should expect a measurable result?

Sixty to ninety days for anything behavioural. Time savings on a single well-defined task can show inside four weeks. Business outcome metrics like win rate need a full sales cycle, so be honest up front about when you can answer, or you will be asked at thirty days and have nothing.

How do we attribute a business outcome to AI specifically?

Cleanly, you often cannot, and claiming otherwise invites an argument you will lose. Say what you can defend: the process metric moved, here is the baseline, here is what else changed in the period. Directional honesty survives scrutiny better than a precise number that does not.

Is a control group worth it?

For a team of eight, no — the overhead exceeds the value. For a rollout across several similar teams, staggering the launch gives you a natural comparison almost for free, and it makes the result substantially more credible.

Key takeaways

  • Baseline for four weeks before you change anything, or the result will be arguable forever.
  • Track outcomes, not activity. Usage rates are an early warning, not evidence.
  • Saved time only counts if it demonstrably went somewhere.
  • Count verification and integration in the cost, not just the subscription.
  • Report at 30, 60, and 90 days — and report the negative results too.