by Claude Opus 5.5

What does a credible 90-day plan look like for moving from experimentation to measurable production value while controlling risk?

A credible plan picks one or two workflows, not ten. It measures a baseline before switching anything on and builds the minimum controls needed to run AI in production. It ships a narrow slice into real work with agreed stop criteria, and ends with a benefits case finance will sign and a repeatable pattern for the next use cases. The test at day 90 is not “how many people are using it” but “what measurably changed, at what cost and risk”.

Why most organisations need this

UK adoption is broad but shallow. ONS data published in July 2026 show about 35% of businesses with 10 or more staff using AI, but the average user firm runs just 1.6 AI technologies, up from 1.4 in late 2023. About 62% of firms citing a lack of AI expertise are training or retraining staff, but only 11% have trained more than half their workforce. Most organisations have trials and licences, not changed processes.

Measurement is the other weak point. The government’s 2024 Copilot experiment across 20,000 civil servants reported 26 minutes saved a day, but the figure was self-reported. When METR measured time directly in 2025, experienced developers were 19% slower with AI while believing they were 20% faster. Self-reports and real effects can point in opposite directions, so the plan has to measure, not survey.

Days 1–30: choose, baseline, set guardrails

Choose one workflow, possibly two. Good first candidates are:

  • high-volume, text-heavy and already measured, such as customer correspondence, internal IT and HR queries, first-pass document review or invoice exception handling;

  • low-harm if wrong, with a human still in the loop;

  • owned by a manager who wants it, rather than one who is being made to have it.

Avoid starting with decisions about people (hiring, performance, discipline). The legal and fairness overhead is real, and it is a poor place to learn.

Measure the baseline now. Take four to eight weeks of current data if you have it: volume, handling time, error or rework rate, escalations, customer outcome, cost per case. If you have nothing, run a two-week manual sample. Without a baseline, day 90 becomes an argument about anecdotes.

Set minimum viable governance:

  • A named business owner accountable for outcomes.

  • A technical owner.

  • Privacy, security and legal involved from week one as design partners, not as a sign-off at the end.

  • A one-page risk assessment and, where personal data is processed in a new way that may be high risk, a data protection impact assessment. The Information Commission (which replaced the ICO on 30 September 2026) expects one for “innovative technologies” likely to result in high risk.

  • An entry in an AI inventory: what the system is, what data it uses, who owns it and which model version it runs.

Agree success and stop criteria in writing. For example: “Handling time down at least 20% with no rise in reopened cases. Stop if the error rate exceeds baseline for two consecutive weeks.”

Days 31–60: build the thin slice and test it

Route usage through a controlled layer. Use enterprise tenants with single sign-on, logging of prompts, outputs and actions, data-loss rules and spending limits. The UK’s Code of Practice for the Cyber Security of AI expects operators to log system and user actions and to keep incident and recovery plans. This is the point to set that up, not after go-live.

Build an evaluation set. Collect 200 to 500 real, anonymised cases with agreed correct answers, written by the people who do the work. Score the system against them before launch and re-run them whenever the model, prompt or data source changes. This is the cheapest insurance available.

Redesign the workflow, not just the tool. Decide where verification sits, what the reviewer checks and what happens to the time saved (see 3.9). Clean up file permissions before connecting any copilot to shared drives.

Talk to staff early. Explain what is being tested and why, and what it means for roles. If your organisation has an information and consultation agreement or a recognised union, check whether the change triggers obligations to inform or consult (see 3.14). Early openness also cuts the shadow AI that Deloitte found among 31% of UK users.

Days 61–90: controlled production and the benefits case

Run a controlled rollout. Use one or two teams with a comparison group, or a staggered start across teams, so you can separate the AI effect from seasonal swings. Build a one-click way for users to flag a wrong or risky output, and a rollback to the previous process.

Track value and risk side by side. On the value side, track time per case, throughput and cost per case. On quality, track the error rate, rework and the share of AI drafts substantially rewritten. On risk, track policy breaches caught, escalations and complaints. On people, track overtime and staff sentiment.

Produce a benefits case finance will sign. It should show baseline against actual results, the full cost stack (licences, usage, integration, review time, governance effort), observed risks and how they were handled, and a decision: scale, adjust or stop. Stopping a use case at day 90 because it failed its criteria is a success of the method, not a failure.

A worked illustration

Take a 300-person UK property-management firm handling 6,000 tenant emails a month, with a measured baseline of 11 minutes per email and 9% reopened. In the first month, the firm baselines performance and assesses the data protection risk, since tenant data is personal. It also decides vulnerable-tenant cases stay human-only. In the second month, AI classifies emails and drafts replies citing the relevant lease clauses, tested against 300 past cases. In the third month, two of six teams go live. By day 90 the comparison shows handling time down to about 7 minutes, with reopen rates flat.

At 4 minutes saved on 2,000 emails a month in the live teams, that is about 130 hours a month. The firm decides to use the time on arrears cases rather than cut staff, and to roll out to the other teams. These numbers are illustrative. The point is that every one of them is measured against a baseline taken before launch.

UK and EU points to check

  • If AI contributes to significant decisions about people, the Data (Use and Access) Act’s safeguards apply: information, a way to contest the decision, and human intervention. The ICO’s draft ADM guidance had still not been finalised by the Information Commission as of early October 2026.

  • In regulated sectors, existing rules apply regardless of technology. In financial services, that includes the PRA’s model-risk expectations (SS1/23) for banks.

  • If you serve EU customers, the EU AI Act’s transparency duties have applied since 2 August 2026. Its high-risk obligations for employment uses now start on 2 December 2027.

Bottom line

A good day-90 outcome is one workflow in production, with numbers measured against a baseline, a full cost line, logs and an evaluation set that can be re-run. The second use case should then take half the time, because the inventory, gateway, test approach and governance route already exist.

Sources

From AI and Jobs: UK, October 2026