by Claude Opus 5.5

How can businesses redesign workflows so AI increases output without increasing cognitive load, rework, or burnout?

Treat checking, not drafting, as the scarce resource. AI makes first drafts almost free, so the bottleneck moves to the people who must judge, correct and take responsibility for what the drafts say. Workflows that stay sustainable decide in advance where verification happens and make it cheap. They stop AI output leaking unreviewed onto colleagues, and they don’t convert every minute saved into more volume.

What the evidence says about where load comes from

The research is consistent on one point: AI helps a lot on some tasks, hurts on others, and people are poor judges of which is which.

  • In the Harvard and BCG field experiment with 758 consultants (Dell’Acqua et al., 2023), those using GPT-4 completed 12.2% more tasks, worked 25.1% faster and produced work rated more than 40% higher in quality. On a task deliberately chosen to sit outside the model’s competence, however, AI users were 19 percentage points less likely to reach the correct answer.

  • In a study of 5,179 customer-support agents (Brynjolfsson, Li and Raymond, published in the Quarterly Journal of Economics in 2025), AI assistance raised issues resolved per hour by 14% on average and by 34% for novice and lower-skilled agents, with minimal gains for the most experienced.

  • METR’s July 2025 trial with 16 experienced open-source developers found they took 19% longer on tasks when allowed to use AI tools. Afterwards they still believed AI had sped them up by about 20%.

  • Researchers from BetterUp Labs and Stanford, writing in Harvard Business Review in September 2025, coined “workslop”: AI output that looks polished but lacks substance, shifting the real work onto whoever receives it. They found 41% of workers had received some, and each instance cost nearly two hours to deal with.

  • A Microsoft Research survey of 319 knowledge workers (CHI 2025) found that higher confidence in AI was associated with less critical thinking, and that effort shifted from producing work to verifying, integrating and stewarding it.

Put together, these studies give the mechanism. AI moves effort from making to checking. It helps most where the checker is a novice being lifted towards a known standard, and least where an expert already works close to their ceiling. When nobody owns the checking, it falls on colleagues downstream. The METR result is the warning for managers: self-reported time savings are not a measure. The UK government’s 2024 Copilot experiment across 20,000 civil servants reported an average saving of 26 minutes a day, but that figure was self-reported, which is why it should be treated as a hypothesis to test rather than a result.

Seven design moves

1. Classify tasks before automating them. For each step, ask four questions. Can the output be checked cheaply, for example against a source document, a calculation or a test? Is the task bounded, with clear inputs and outputs? Can it be reversed if it is wrong? Who is harmed if it is wrong? Use AI heavily where outputs are cheap to check and the stakes are low. Keep it advisory where checking is expensive and the stakes are high. The worst combination is “expensive to check, but looks right”, which is where the 19-point accuracy penalty lives.

2. Put verification in one named place. Rework multiplies when everyone half-checks. Decide which role signs off each output and what “checked” means: sources opened, figures reconciled, tone confirmed. Then remove duplicate review elsewhere. One careful review beats three cursory ones.

3. Make checking cheaper than generating. Require outputs to cite the passages they rely on, so a reviewer can verify a claim in seconds. Show changes as tracked edits rather than full rewrites. Have automated checks flag unsourced claims, unreconciled numbers and personal data before a human sees the draft. If reviewing an AI draft takes as long as writing it, the workflow has failed.

4. Replace blank prompts with structured requests. Open-ended prompting is itself cognitive work. Pre-built actions with fixed fields (audience, policy reference, required sources) and fixed output templates reduce variation, and it is variation that creates rework.

5. Adopt a sender-owns-it rule. Anyone who sends AI-assisted content to a colleague or client is responsible for it as if they had written it unaided. This single norm does more against workslop than any tool setting. It also makes AI use discussable rather than hidden, which matters when Deloitte’s 2026 survey of 25,000 UK workers found 31% of GenAI users using it without their employer knowing.

6. Bank part of the time saved. If every saved minute becomes extra quota, intensity rises and errors follow. Decide explicitly how gains are split between more throughput, better quality, training and slack for exceptions. Track work intensity (queue length, after-hours activity, sickness absence) alongside output.

7. Protect learning for juniors. Since AI lifts novices most, it is tempting to let them lean on it entirely. Keep some work done without AI and reviewed by an expert, so that the people who will be the checkers in five years learn what good looks like (see 3.10).

A worked example

Take a 25-person customer-correspondence team at a UK insurer handling 4,000 complaint letters a month. The naive rollout gives everyone a chat assistant. Each handler then generates drafts, half-checks them, and team leaders re-check anything sensitive. Output rises a little, rework rises more, and team leaders become the bottleneck.

The redesigned version works differently:

  • AI classifies each letter by complaint type and risk. It routes vulnerable-customer and regulatory-deadline cases straight to experienced handlers.

  • For routine cases, AI drafts a reply in a fixed template, quotes the policy clauses it relied on, and flags any figure it could not reconcile.

  • The handler reviews the tracked draft against those clauses and signs it off as the single point of verification. Team leaders sample 5% of sign-offs rather than re-checking everything.

  • Two hours a week per handler are protected for complex cases and coaching.

The metrics are first-time resolution, complaints reopened, the share of drafts substantially rewritten, and overtime. Speed alone is not the measure. If the rewrite rate stays high for a complaint type, that type comes out of the AI route.

The UK duty-of-care angle

This is not only good practice. HSE states that employers “have a legal duty to protect employees from stress at work by doing a risk assessment and acting on it”. Its Management Standards cover six areas: demands, control, support, relationships, role and change. An AI rollout touches at least three of them. It changes demands (volume targets), control (how much say people have over method) and change (how the transition is managed and communicated). Running the HSE risk assessment for a major AI workflow change is a practical way to catch burnout risks before they show up in absence data.

One caution: it is tempting to use AI usage logs to monitor individual productivity. The ICO’s guidance on monitoring workers (now under review following the Data (Use and Access) Act) warns that “just because a form of monitoring is available, does not mean it is the best way to achieve your aims”. It also says that monitoring “conducted without transparency is unfair”, and it lists monitoring that may lead to financial loss, such as through performance management, among the uses likely to need a data protection impact assessment. Heavy-handed monitoring also erodes the trust that honest reporting of AI errors depends on.

Bottom line

Ask four questions of any AI workflow change. Did we remove steps, or just add a tool? Is checking now cheaper than producing? Is there one clear owner of verification? Have we decided where the time saved goes? If any answer is no, the organisation is likely buying output with its staff’s attention, and that bill arrives later as rework, errors and turnover.

Sources

From AI and Jobs: UK, October 2026