by Claude Opus 5.5
How do agentic AI systems—which carry out multi-step tasks with limited supervision—change the jobs/tasks/workflows framework, and do they shift the balance from augmentation towards substitution?
Agents move the unit of automation from the task to the workflow. A copilot speeds up a step that a person still performs. An agent carries out a sequence of steps itself, using software tools, and the person’s role shifts from doing the work to specifying it, reviewing it and handling exceptions. That does tilt the balance towards substitution, especially in digital, rule-bound, high-volume work whose results can be checked cheaply. In 2026, though, reliability, compounding errors, security and accountability keep most agent deployments supervised. Robust data on how many UK firms use agents barely exist.
What an agent does differently
Three features separate an agent from a chatbot or copilot.
Multi-step execution. The system breaks a goal into steps, carries them out, checks intermediate results and adjusts. A copilot answers a prompt and stops.
Tool use. The agent acts through software: it searches, reads files, queries databases, fills in forms, sends messages and runs code. That lets it do the “glue” work between systems that used to need a person.
Delegated initiative. The person sets a goal and constraints rather than giving each instruction, and the agent decides the order of work within those limits.
Take a supplier invoice that doesn’t match its purchase order. A copilot helps a clerk write a query email. An agent reads the invoice, finds the purchase order in the finance system, checks the goods-received record, identifies a price discrepancy, emails the supplier, posts the invoice when a credit note arrives, and escalates to a human only if the supplier disputes it. In the first case the clerk does the job faster. In the second the clerk becomes the exception handler for many such cases.
How the jobs/tasks/workflows framework changes
The standard framework treats a job as a bundle of tasks. AI exposure is usually measured task by task, asking whether a model could do this activity. That approach understates agents, for two reasons.
First, agents attack the hand-offs, not just the tasks. Much white-collar work consists of moving information between systems, chasing people, reconciling records and assembling outputs. These activities rarely appear as distinct tasks in occupational databases, but they take up a large share of junior and administrative time. Removing them can remove a role even when no single listed task looked highly exposed.
Second, the human role changes in kind, not just in speed. The difference is easiest to see side by side:
Who initiates each step. Copilot (augmentation): Human. Agent (delegation): Agent, within a brief.
Unit of work. Copilot (augmentation): A task. Agent (delegation): A workflow or case.
Human role. Copilot (augmentation): Doing, faster. Agent (delegation): Specifying, reviewing, handling exceptions.
Typical failure. Copilot (augmentation): Poor draft, caught by the author. Agent (delegation): Wrong action, possibly unnoticed.
Where labour is saved. Copilot (augmentation): Time within a role. Agent (delegation): Number of people per unit of volume.
The skills that gain value are those needed to brief, check and intervene: knowing the process well enough to define it, judging whether an output is right, and handling the cases that fall outside the rules. Those happen to be skills people have traditionally learned by doing the routine work that agents now take over. That is the entry-level problem discussed in answer 1.8, in sharper form.
A shift towards substitution, within limits
Yes, in a specific and bounded way. The best usage evidence comes from how firms deploy models programmatically, which is where agents run. Anthropic’s Economic Index (September 2025) found that 77% of enterprise API transcripts showed automation patterns, mostly full task delegation, against 12% augmentation. Consumer use of Claude.ai was roughly evenly split. Firms that wire AI into their systems are using it to do work, not to help someone do it.
Substitution pressure rises most where five conditions hold:
The workflow is fully digital from end to end.
Correct outputs can be checked cheaply, for example by reconciling against a record or running a test.
Errors are reversible or low-cost.
Volumes are high, so building and supervising the agent pays off.
The rules are stable and documented.
Back-office finance, claims and case processing, first-line customer service resolution, data entry and reconciliation, routine code maintenance and recruitment administration fit this profile. Bank of England staff analysis on the Bank Underground blog (August 2026) found customer-service and administrative vacancies down more than 20% over three years. That is consistent with this pattern, but it does not show that agents caused it, since weak demand and cost pressures have been strong over the same period.
The labour arithmetic turns on review cost. Here is an illustration with assumed numbers. A ten-person team handles 1,000 cases a week. An agent resolves 700 cases end to end, and humans check them at an average of 20% of the time a full case takes, while handling the other 300 cases themselves. Human workload falls to the equivalent of 440 full cases (300 + 700 × 0.2), about 4.4 people. If thorough review instead takes 50% of a case’s time, workload is 650 cases, or 6.5 people. The same agent delivers a 56% or a 35% reduction depending almost entirely on how much checking it needs. That is why reliability, not raw capability, governs substitution.
Substitution is also more likely to show up as slower hiring and smaller junior intakes than as redundancy. A team that shrinks from ten people to six usually gets there by not replacing leavers.
The realistic limits in 2026
Reliability. METR’s 2025 research measures the length of tasks, timed by how long human professionals take, that AI agents can complete autonomously. It found this “time horizon” had been doubling roughly every seven months for six years. But the headline measure is the task length at which agents succeed 50% of the time, and the best model in March 2025 managed about one hour. Business processes need far higher success rates than 50%. Models released in mid-2026, including OpenAI’s “ChatGPT Work” agent and Anthropic’s Claude Opus 5, are more capable. No independent evaluation of their reliability in UK workplace conditions has been published. On Carnegie Mellon’s TheAgentCompany benchmark, a simulated software company, the best agent completed 30% of tasks autonomously in the September 2025 revision.
Error compounding. Errors multiply across steps. As simple arithmetic: if each step of a 20-step process succeeds 95% of the time and errors are independent, the whole process succeeds only about 36% of the time (0.95²⁰). At 99% per step it succeeds about 82% of the time. Long workflows therefore need either near-perfect steps or checkpoints, and checkpoints bring humans back in. Errors can also be silent. A wrong posting or a misfiled document may not surface until an audit or a complaint.
Security. Agents read untrusted content, such as emails, web pages and supplier documents, and can also act. The NCSC warned in December 2025 that prompt injection, where hidden instructions are planted in content an AI reads, may “never be properly mitigated” because language models do not reliably separate data from instructions. An agent with permission to pay invoices is a target in a way a drafting assistant is not. Least-privilege access and human approval for irreversible actions are the standard responses, and both reduce the labour saving.
Accountability. UK law does not let an organisation blame its software. Under the Data (Use and Access) Act’s framework, in force since 5 February 2026, significant automated decisions about individuals require safeguards, including the right to contest and to obtain human intervention. In financial services the Senior Managers and Certification Regime makes named individuals answerable, and industry has questioned how the PRA’s model-risk expectations (SS1/23) scale to agentic systems. Someone has to own each agent’s actions, and owning them means monitoring them.
The verification trap. Checking AI output is real work, and people misjudge it. In METR’s 2025 trial, experienced developers took 19% longer with AI tools while believing they had been about 20% faster. Reviewing an agent’s case file can take nearly as long as doing the case if the reviewer can’t trust its intermediate steps.
Hype. Gartner predicted in June 2025 that over 40% of agentic AI projects would be cancelled by the end of 2027, citing cost, unclear value and weak risk controls. It also warned of “agent washing”: by its estimate only about 130 of thousands of vendors offered genuinely agentic products.
What we know about UK adoption
Very little, robustly. The best official figure is DSIT’s finding that 7% of AI-adopting businesses used agentic AI, from fieldwork in February to May 2025. Since only about one in six businesses used AI at all in that survey, agents were a small minority practice. The ONS Business Insights survey does not ask about agents. Vendor surveys claiming high agent adoption among UK firms rarely publish sample details or definitions, and given Gartner’s warning about agent washing, a firm reporting “agents” may mean a scripted chatbot. Treat any precise UK agentic adoption figure for 2026 with caution.
Bottom line
Agents make substitution more plausible than copilots did, because they target whole workflows and the connective work between tasks, much of which is junior work. In 2026 the binding limit is the cost of making them reliable and accountable, not their capability in demonstrations. Over the next few years the labour effect will depend less on what agents can do than on how cheaply their work can be verified.
What to watch
Independent reliability results at the 80–95% success thresholds that businesses need, not 50%.
Results of the Bank of England/FCA AI survey due at the end of 2026, the first large UK dataset likely to capture agentic use in finance.
Whether insurers and regulators set expectations for agent permissions and audit trails.
Ratios of reviewers to case volumes in back-office functions, and the size of junior intakes.
Sources
Anthropic Economic Index report, September 2025 — Anthropic, Sep 2025
Measuring AI ability to complete long tasks — METR, 19 Mar 2025
UK data protection and privacy reform goes live — HSF Kramer, 5 Feb 2026
UK financial services regulators’ approach to AI in 2026 — Covington, 9 Apr 2026
Keeping an eye on AI in financial services: the AI survey — Burges Salmon, 22 Jul 2026