by Claude Opus 5.5
How should individuals work effectively with AI agents—what to delegate, how to verify outputs, and who is accountable when an agent gets it wrong?
Delegate tasks whose results you can check more cheaply than you could produce them yourself. Brief the agent as you would a capable contractor who knows nothing about your organisation. Build in stopping points before anything that cannot be undone. Check the work against independent evidence, not against the agent’s own account of it. Give it the minimum access it needs. And assume you are answerable for what it does: your employer will hold you to account, and professional regulators say the same.
What makes agents different
A chatbot drafts and you decide what happens next. An agent acts. It runs searches, edits files, sends messages, updates records and executes code, across many steps with little supervision. That changes the risks in three ways:
Errors compound. A wrong assumption at step 2 can carry through step 20 unnoticed.
Actions can be irreversible. An email sent, a record deleted or an order placed cannot be undone by editing a draft.
The agent reads untrusted content. A web page or document can contain hidden instructions. OWASP ranks this “prompt injection” as the top security risk for large language model applications, including “indirect” injection from external websites or files.
Agent products moved into the mainstream in mid-2026, including OpenAI’s “ChatGPT Work” agent in July, but UK workplace use is still early. DSIT’s adoption research, based on fieldwork in 2025, found only 7% of firms using AI were using agents. Many people will meet agents first through features built into tools they already use, so it is worth building good habits now.
What to delegate
Six questions help:
Is it cheap to check? If verifying the result takes almost as long as doing the work, there is little to gain, and real risk in skipping the check.
Can it be undone? Reversible tasks such as drafts, analyses or a branch of code are good candidates. Irreversible ones need approval steps.
How far could a mistake spread? An error in your own notes is contained. An error in a client communication or a shared database is not.
Is “done” well defined? “Reconcile these two spreadsheets and list the mismatches” works. “Sort out the supplier situation” does not.
Is it inside the agent’s competence? In the Dell’Acqua experiment with consultants, AI raised quality by more than 40% on tasks within its capability. Outside that “jagged frontier”, consultants using AI were 19 percentage points less likely to reach the right answer. Ability can drop sharply on tasks that look similar.
What data does it touch? Personal, confidential or client data needs an approved tool and a clear lawful basis, whatever the agent’s capability.
Research summaries with links, data cleaning, first-draft analyses, test generation, reformatting, triaging your own inbox into folders. Delegate with checkpoints: Multi-file code changes, updating internal records, drafting replies to external contacts, booking within set rules. Keep for yourself, or approve every step: Payments, contractual commitments, decisions about individuals, external publication, deleting data, anything regulated.
Writing the brief
Most agent failures trace back to an unclear brief. A good one covers:
Goal and purpose: what you want and why. The “why” lets the agent make sensible small decisions.
Context: the facts it cannot infer, such as who the audience is, what has been tried and which conventions apply.
Inputs and sources: which files, systems and websites it may use, and which it may not.
Constraints: budget, deadline, style, and actions it must not take.
Definition of done: the output format and the tests it must pass.
When unsure: “Stop and ask” beats “make a reasonable assumption” for anything consequential.
Evidence: ask it to return links, file paths, the queries it ran and a list of the assumptions it made.
Checkpoints
For anything beyond a short, reversible task, structure the work:
Plan first. Ask for a plan and approve it before execution. Most misunderstandings surface here at almost no cost.
Gate irreversible actions. Require approval before anything is sent, paid, deleted, published or merged, using the tool’s settings where possible.
Limit the run. Cap steps, time or spending, so a confused agent stops rather than wandering.
Check intermediate work. For long tasks, review the interim outputs (the extracted data, the draft structure) before the agent builds on them.
Verification techniques
The essential rule is to check the work, not the agent’s report of the work. Agents can report a task as complete when it isn’t, or describe a step they did not take.
Check the action happened. Open the file, look in the sent folder, query the record. Don’t rely on the summary.
Go back to the source. For research, open the links and confirm they say what is claimed. For summaries, compare against the original for omissions, not just errors.
Recompute the key numbers independently, even roughly.
Read the diff. For changes to code or documents, review exactly what changed, line by line for anything that matters.
Test edge cases. Run cases where you know the right answer, including awkward ones.
Sample in proportion to risk. For batch work such as 500 categorised invoices, check a random sample plus every high-value or unusual item.
Watch for scope creep. Did it touch anything outside the brief? Check logs where available.
Permissions hygiene
Least privilege. Grant read-only access first. Add write access for specific folders or systems only when the task needs it. OWASP’s guidance is to restrict a model’s privileges “to the minimum necessary” and to require human approval for privileged operations.
Use scoped, temporary access. Prefer task-specific connections that expire over standing access to your whole email, drive and calendar.
Don’t combine dangerous capabilities. An agent that can read untrusted web content and send email and access confidential files is a data-leak waiting to happen. Split these across tasks.
Use approved tools only. Deloitte found that 31% of UK workers using generative AI do so without their employer knowing. With agents, an unapproved tool can act inside company systems, not just read from them.
Review and revoke. Periodically check which agents and integrations hold access to your accounts, and remove those you no longer use.
Accountability
To your employer. In practice, an agent acting with your credentials, at your instruction, on work you submit, is treated as your tool. If it emails the wrong client or corrupts a dataset, you will be asked what you delegated, what controls you set and what you checked. Responsibility is shared to the extent that your employer approved the tool and defined the process, which is a good reason to ask for that definition in writing if it does not exist. Keep a brief record of what you delegated and how you verified it. Report errors promptly; quick disclosure is almost always treated better than discovery.
To professional regulators. The bodies that have spoken are consistent. The Solicitors Regulation Authority tells firms they “must remain responsible” for their activities and “cannot delegate accountability to an IT team or external provider”. It gives the example that a firm using a third-party chatbot for initial legal advice remains “responsible for any errors”. The General Medical Council says doctors “are responsible for the decisions they take when using new technologies like AI”. The courts have taken the same line. In Ayinde v Haringey (2025) the Divisional Court held that lawyers must check AI-assisted research against authoritative sources, and likened this to the duty to supervise work done by others. If you belong to a professional body, read its current AI guidance. Expect it to treat an agent like a junior colleague whose work you supervise, not as a separate party that shares the blame.
To third parties. Customers will generally hold the organisation responsible, and the organisation will look to the people who ran the process. “The agent did it” is not a defence to plan on. For specific legal exposure, take advice.
Bottom line
Treat an agent like a fast, tireless contractor who has never worked for your organisation. Give it a clear brief, limited access and approval points before anything irreversible, and inspect what it delivers. The productivity gain is real when verification is cheaper than doing the work. The accountability stays with you.