by Claude Opus 5.5

What’s a practical framework a business can use to decompose roles into tasks and classify each task as AI-automatable, AI-augmentable, or human-critical?

List the 12–25 tasks that actually fill a role and record the share of time each takes. Score every task on five tests: how cheaply the output can be checked, what an error costs, how often cases fall outside the rules, whether a named human must be accountable, and whether AI can lawfully reach the data. Then classify the task within its workflow, not the job title. A role rarely lands in one bucket. In the illustrative accounts-payable example below, about 60% of time is automatable; in the paralegal example, far less is. In both, the small human-critical residue is where most of the risk sits.

Step 1: build an honest task inventory

Job descriptions are a poor source because they describe the job as advertised, not as done. Use three inputs instead:

  • A starting list. Skills England’s occupational standards list the duties for hundreds of roles. The Paralegal standard, for example, sets out 14 duties, from “complete routine legal research” to “identify personal competence limits”.

  • A two-week time diary from three or four post-holders, coded to that list.

  • System logs, which show volumes, exception queues and hand-offs.

For each task, record its share of time, frequency, inputs, outputs, the system it runs in and who receives the result. Twelve to 25 tasks is the right grain. Fewer hides the variation; more turns the exercise into time-and-motion study.

Step 2: score each task on five tests

  • Verifiability. Question to ask: Can a competent person check the output much faster than produce it? Pushes towards: Automate or augment if yes.

  • Error cost. Question to ask: If it is wrong and nobody notices for a week, what happens, and is it reversible? Pushes towards: Human-critical if severe and irreversible.

  • Exception rate. Question to ask: What share of cases need context that isn’t in the system (a phone call, local knowledge, a judgement about intent)? Pushes towards: Augment if high.

  • Accountability. Question to ask: Must a named person carry the decision by law, regulation or professional rule? Pushes towards: Human-critical.

  • Data access. Question to ask: Can a tool reach the inputs, and is it lawful and secure for it to do so? Pushes towards: Blocks automation until fixed.

Score each test 1–3 and write one line of justification. The justification matters more than the number, because it is what you will revisit when tools improve.

The accountability test has a new UK edge. Since the Data (Use and Access) Act 2025’s automated decision-making provisions commenced on 5 February 2026, solely automated significant decisions about individuals are permitted, but only with safeguards: people must be told, must be able to contest the decision and must be able to get human intervention. Special category data remains restricted. A task that produces a significant decision about a person can therefore move towards automation, but only if you build those safeguards. The ICO’s ADM guidance is still in draft, and the Information Commission (which replaced the ICO on 30 September 2026) has not yet finalised it.

Step 3: classify with explicit rules

  • AI-automatable. The output is cheap to verify by sampling, errors are low-cost or reversible, the exception rate is low and no named accountability attaches. “Automatable” does not mean unattended. It means AI handles the normal case end to end, low-confidence cases are routed to a person, and a sample is checked.

  • AI-augmentable. AI produces a draft, an analysis or a set of options, and a human decides. This fits tasks where drafting takes much of the time but errors are costly or the context is rich.

  • Human-critical. A named person must be accountable, the value depends on a relationship or on trust, or checking the AI’s work would take as long as doing it.

Agentic tools change how you apply the rules. A chain of tasks that are each only augmentable can become automatable when an agent can carry context from one step to the next. Errors also compound along that chain. So classify each hand-off as well as each task.

Worked example 1: a UK accounts-payable clerk

The time shares below are illustrative, for a clerk in a mid-sized firm processing several thousand invoices a month.

  • Capture invoice data from PDFs and emails. Time: 25%. Class: Automatable. Reasoning: Extraction with confidence thresholds; low-confidence fields go to a person. Mandatory e-invoicing for VAT invoices from 2029 will shrink this task further.

  • Three-way match (PO, goods receipt, invoice). Time: 20%. Class: Automatable for clean matches. Reasoning: Rules plus AI for fuzzy line matching.

  • Resolve match exceptions (price variances, part deliveries). Time: 15%. Class: Augmentable. Reasoning: AI gathers the evidence and proposes a fix; the context sits with buyers and warehouses.

  • Answer supplier “where’s my payment?” queries. Time: 10%. Class: Mostly automatable. Reasoning: Status lookups are routine; disputes are escalated.

  • Prepare the payment run. Time: 5%. Class: Automatable. Reasoning: Prepared by system.

  • Approve and release payments. Time: 5%. Class: Human-critical. Reasoning: Authorisation limits and segregation of duties.

  • Change supplier master data, including bank details. Time: 5%. Class: Human-critical. Reasoning: This is where mandate fraud happens. Verify by call-back to a known number.

  • Flag duplicates and anomalies. Time: 5%. Class: Augmentable. Reasoning: AI flags; a person investigates.

  • Month-end accruals and supplier statement reconciliations. Time: 10%. Class: Augmentable. Reasoning: Judgement on timing and disputes.

On these assumptions, about 55–60% of time sits in automatable tasks, about 30% in augmentable ones and about 10% in human-critical ones.

That does not mean that 60% of AP clerks go. Automated work still needs exception handling and sample checks, and the human-critical 10% grows in importance. Large organisations have faced the “failure to prevent fraud” offence since 1 September 2025, and their main defence is having reasonable prevention procedures. Their AP teams therefore cannot treat bank-detail changes as admin. An AI agent that reads supplier emails is also an obvious target for instructions planted in an invoice.

A plausible redesign turns a team of six clerks into a smaller team of exception handlers and controls owners. Whether it shrinks depends on volumes and on what the freed time is redeployed to. One cost is easy to miss: data capture and matching were how junior staff learned the ledger, so the training route has to be rebuilt deliberately.

Worked example 2: a paralegal

Applying the same tests to the Skills England paralegal duties gives a different profile.

  • Routine legal research. Augmentable, never automatable. In Ayinde v Haringey (June 2025) the Divisional Court said those using AI for research have a professional duty to check it against authoritative sources. Two of the lawyers before it had relied on non-existent cases: five in one matter and 18 in the other.

  • Initial document review (disclosure, contracts). Augmentable. AI does the first pass, and a person samples the results and makes privilege calls.

  • First drafts of letters and documents. Augmentable, with a supervising lawyer’s sign-off.

  • Bundles, chronologies and file administration. Largely automatable.

  • KYC and anti-money-laundering checks. Split. Collecting documents and screening are automatable; the risk judgement remains human-critical.

  • Client contact on routine matters and escalating beyond one’s competence. Human-critical. The second is human-critical by definition.

The paralegal’s profile shifts further towards augmentation than the AP clerk’s, because the cost of a missed error falls on a client and a court. Narrow automated legal services do exist. The SRA authorised Garfield.Law, an AI-driven firm handling small debt claims, in May 2025. But they work within tight scopes and keep solicitors accountable.

Common mistakes

  • Classifying the job instead of the task. “Paralegals are exposed” tells you nothing you can act on.

  • Ignoring checking time. A task is only worth augmenting if verifying the output is genuinely cheaper than producing it. In a BCG field experiment, consultants given AI on a task outside its capability were 19 percentage points less likely to reach the correct answer than those working without it.

  • Treating the scores as permanent. Capability changes quickly; Anthropic, OpenAI and Google all shipped major models in summer 2026. Re-score priority roles every six months.

  • Forgetting that tasks also train people. For each task you classify as automatable, note what a junior learned from doing it, and decide where that learning will now happen.

Bottom line

The framework’s value lies less in the three labels than in the written reasoning behind each score. That reasoning shows where a role’s risk and judgement are concentrated, which tasks protect the business, and what has to be rebuilt for the next generation of staff.

Sources

From AI and Jobs: UK, October 2026