by Claude Opus 5.5

What metrics best capture real AI impact beyond usage, and how should baselines be set?

Real impact is the change in a workflow’s outcomes compared with what would have happened without AI, net of AI’s full cost. That means quality-adjusted output, unit cost, cycle time and the customer outcome, watched alongside risk and workforce measures. Baselines should be taken before launch, cover a full business cycle and, where possible, use a comparison group that has not yet received the tool. Usage figures and self-reported time saved are the weakest evidence available. UK government trials have shown how far they can diverge from measured productivity.

Why usage and “time saved” mislead

Three pieces of evidence make the point.

  • Self-reports and productivity can diverge. In the cross-government Microsoft 365 Copilot experiment, 20,000 civil servants reported saving an average of 26 minutes a day. Those figures were self-reported, using time bands. The Department for Business and Trade’s separate evaluation “did not find evidence that time savings have led to improved productivity”.

  • People misjudge their own speed-up. METR’s 2025 trial found experienced open-source developers were 19% slower with AI tools on their own codebases. Afterwards they believed AI had made them 20% faster.

  • Averages hide who gains. In Brynjolfsson, Li and Raymond’s study of 5,179 support agents, productivity rose 14% on average and 34% for novices, with little effect for the most experienced. An average-only dashboard would miss where the value sits, and the risk of de-skilling that comes with it.

Usage matters, but it is an input, like electricity consumption, not an outcome.

The metric stack

Pick one primary outcome and protect it with guardrails. The examples below cover an accounts-payable team and a contact centre.

  • Quality-adjusted output (primary). What it measures: Units of finished work per paid hour, net of rework. Accounts payable: Invoices fully processed per FTE-hour, excluding those reopened. Contact centre: Contacts resolved per agent-hour, excluding repeat contacts within 7 days.

  • Unit cost. What it measures: All-in cost per unit, including licences, usage charges, review time and support. Accounts payable: Cost per invoice. Contact centre: Cost per resolved contact.

  • Cycle time. What it measures: Elapsed time end to end, not time on task. Accounts payable: Receipt to approval-ready, in days. Contact centre: Time to resolution.

  • Quality and errors. What it measures: Errors that escape, found by blind sampling. Accounts payable: Duplicate or incorrect payments; late payments against terms. Contact centre: Complaint rate; QA-scored accuracy.

  • Customer outcome. What it measures: What the recipient experiences. Accounts payable: Supplier queries per 1,000 invoices. Contact centre: First-contact resolution; customer satisfaction.

  • Risk guardrails. What it measures: Signals that something is going wrong. Accounts payable: Override rate on AI proposals; bank-detail change exceptions. Contact centre: Escalations; vulnerable-customer handling failures.

  • People. What it measures: Effects on the workforce. Accounts payable: Time to competence for new starters; overtime; attrition. Contact centre: Same, segmented by tenure.

  • Capacity use. What it measures: Where freed time went. Accounts payable: Backlog cleared; new work taken on; vacancies not refilled. Contact centre: Same.

Two layers deserve emphasis. Capacity use turns minutes into money or service. If it shows nothing, the saving is not real. Vacancies not refilled belongs with the people measures because that is how AI’s headcount effect mostly shows up. The Bank of England’s July 2026 Monetary Policy Report said AI adoption is reducing demand for highly automatable jobs in some industries, “with firms often slowing hiring or leaving vacancies unfilled”. A firm that tracks only redundancies will under-count its own workforce effects.

Setting baselines

1. Define the unit and the boundary. Measure per invoice, claim, contact or matter, and include downstream rework and escalations. AI often moves effort around instead of removing it.

2. Measure for long enough. Take 8 to 12 weeks before launch at minimum, and preferably a full cycle. For accounts payable that means including a month-end and, ideally, the year-end peak. Record case mix as you go: invoice types, supplier counts and contact reasons. Otherwise a quieter quarter will look like an AI effect.

3. Choose the strongest comparison you can manage. In descending order of credibility:

  • Randomised access by team or queue. This is feasible for copilots and agent-assist tools, and is the cleanest test.

  • Staggered roll-out. Teams not yet switched on act as controls (a difference-in-differences design). This is usually the most practical choice, since roll-outs are staggered anyway, and it is the design behind the customer-support evidence above.

  • Matched comparison with a similar site, entity or team.

  • Before and after with adjustments for seasonality and volume. This is the weakest design; label it as such.

4. Exclude the ramp-up. Report the first 4 to 6 weeks separately. Results usually dip while people learn, then recover. Check again at six months to see whether gains last beyond the novelty.

5. Separate the tool from the redesign. If you change routing, approval steps or roles at the same time, label the result “AI plus process change”. Do not credit it all to the model; the redesign may be worth keeping even if the tool is not.

6. Score quality blind. Use a fixed rubric, raters who do not know which outputs had AI help, and a separate count of critical errors. One wrong payment or fabricated citation can outweigh a hundred faster drafts.

7. Keep a small hold-out. A team or queue left without AI for a further quarter lets you keep checking that the effect is real, and shows when model updates help or hurt.

Worked example (illustrative figures)

A UK distributor’s AP team has 6 FTE at a fully loaded £38,000 each, which comes to £19,000 a month. It processes 9,000 invoices a month, so its baseline cost is £2.11 per invoice. In its baseline quarter, 14% of invoices were paid late against terms and duplicate payments were found in about 1 in 2,000 invoices.

After launching invoice extraction and matching, with a three-month steady-state period following a six-week ramp-up:

  • 55% of invoices pass through untouched, and clerks handle exceptions.

  • The work now takes 4.2 FTE, costing £13,300 a month. The tool costs £3,000 a month. The all-in cost is £16,300, or £1.81 per invoice, a fall of about 14%.

  • A sister site not yet switched on saw its cost per invoice fall about 4% over the same period, as volumes rose. The effect attributable to AI is therefore roughly 10%. That is well short of the vendor’s projection, mainly because exceptions and review take real time.

  • Late payments fell to 9%. Duplicate payments did not change, which is a signal to look again at the anomaly checks.

  • The 1.8 FTE freed were moved to supplier statement reconciliation. That counts as value only because the team can show recovered credits and fewer supplier queries. Had the time gone nowhere measurable, there would be no real saving, only the tool’s cost.

This is what an honest result looks like: positive, smaller than the sales pitch, with one guardrail flagging work still to do.

Pitfalls

  • Goodhart’s law. Reward “invoices processed” alone and quality will suffer, so always pair throughput with an escaped-error measure.

  • Double counting. The same saved hour should not appear in both the AI programme’s benefits and a separate efficiency programme’s.

  • Survivorship. If struggling users drop the tool, the remaining users’ figures flatter it. Report effects for everyone who was given access, not just those who kept using it.

  • Headcount as the only result. About 6% of UK firms using AI in their operations report a fall in headcount (ONS, July 2026). Most gains show up as capacity, quality or avoided hiring, which only the fuller stack captures.

Bottom line

Measure the workflow, not the tool. Baseline it before launch, compare against teams that do not have the tool, and count only the time saved that you can see in output, cost or quality.

Sources

From AI and Jobs: UK, October 2026