by Claude Opus 5.5

What kinds of portfolio evidence best demonstrate real AI leverage?

The strongest evidence is a before-and-after case on real work. It shows the task, the baseline, what you built or changed, how you checked the output, what went wrong and the measured result. Screenshots of prompts, lists of tools and stacks of course certificates are weak, because they show exposure rather than leverage. One well-documented case where your judgement is visible beats ten demos.

Why portfolio evidence matters more in 2026

Employers are wading through volume and struggling to read signals. The Institute of Student Employers’ 2025 survey found large employers receiving 140 applications per graduate vacancy. In the same survey, 61% had seen candidates use AI in interviews without disclosing it. When everyone’s CV and cover letter are AI-polished, claims cost nothing, and verifiable work samples become the signal.

The type of evidence matters as well. DSIT and LinkedIn’s June 2026 entry-level snapshot found employers seeking “specific, operational, production-ready capabilities”. Candidates were offering “general, analytical skills”. A portfolio is the most direct way to show you are on the production-ready side of that line. With only 21% of UK workers saying they feel confident using AI at work (government research, January 2026), credible proof of applied skill still stands out.

What makes AI evidence credible

Before choosing artefacts, check that each one has six properties:

  1. A real task. Use work that someone actually needed done, or a faithful synthetic version of it, rather than a toy problem.

  2. A baseline. Record how long it took, or how good it was, before you changed anything.

  3. A verification method. Explain how you knew the AI output was right: sampling, reconciliation, test cases, expert review.

  4. A failure log. Show what the tool got wrong and what you did about it. This is the single most persuasive element, because it proves you were checking.

  5. Visible judgement. Point to where you overrode, constrained or redirected the tool.

  6. Safe handling. Use no real client, patient or employee data. Anonymise, use synthetic data, or describe the work without exposing it.

Seven artefacts that work

1. The one-page workflow case study. This is the core format. Set it out under headings: Problem, Baseline, What I built, How I checked it, Result, What I’d change. Keep it to a page and include one diagram of the workflow. Two or three of these, from different parts of your job, make a strong portfolio.

2. An evaluation set. This is a test pack used to judge an AI tool on your kind of work. For example, a paralegal might assemble 40 anonymised lease clauses with the correct reading of each, run them through the firm’s summariser, and record which clause types it misread. Few people produce one, and it demonstrates verification directly rather than asserting it.

3. A reusable asset others use. Examples include a configured assistant, a template library with version notes, or a checking checklist. Adoption is the evidence. “Used by the 12-person claims team since March” says more than any description of features.

4. A small automation. Examples include a Power Automate or Zapier flow, a spreadsheet with AI-extracted fields and validation rules, or a short Python script. Show the error handling, not just the happy path.

5. A decision record. This is a short memo showing a case where AI analysis informed a decision and you added the judgement. It should cover what the model suggested, what you checked, and what you decided and why. It works well for managers and professionals whose output is decisions rather than artefacts.

6. A teaching artefact. Examples include a guidance note your team adopted, slides from a session you ran, or a “how we use AI for X” page with review rules. These show the senior skills that PwC’s US analysis finds AI-exposed junior roles increasingly demand.

7. Public work, where the field expects it. For developers, analysts and designers, a GitHub repository, a notebook or a short write-up counts. It should explain how AI was used and how the result was tested. Repositories that simply contain generated code with no tests prove little.

Examples by role

  • Customer service adviser. Strong artefact: Verified response library with a log of bot answers you corrected and why. What it proves: Quality control; knowledge of policy and vulnerability handling.

  • Claims handler. Strong artefact: Triage checklist for AI-flagged claims, with a sample showing false positives and missed fraud indicators. What it proves: Exception judgement; fraud awareness.

  • Accounts payable. Strong artefact: Exception rules for AI-matched invoices, with error rates by supplier before and after. What it proves: Process design; controls thinking.

  • Paralegal. Strong artefact: Evaluation set for a contract-review tool, plus a review playbook. What it proves: Verification; legal knowledge applied.

  • Marketing executive. Strong artefact: Content workflow showing brand and legal checks, with performance compared against the previous approach. What it proves: Brand judgement; measurement.

  • HR adviser. Strong artefact: Policy assistant tested against real anonymised queries, with escalation rules for employee relations cases. What it proves: Safe deployment; knowing what not to automate.

  • Teacher. Strong artefact: Adapted AI-drafted resources, showing the changes made for a specific class and the errors caught. What it proves: Pedagogical judgement.

  • Junior developer. Strong artefact: Feature built with an assistant, with tests, a security review and notes on rejected suggestions. What it proves: Engineering discipline.

Writing the numbers honestly

Report what you measured and how. For example: “Average time to first draft of a complaint response fell from about 25 to about 10 minutes, measured over 40 cases; one in eight drafts needed a material correction.” Those numbers illustrate the format only. If you did not measure, say so and describe the qualitative change. Inflated or unmeasured percentages are easy to spot in interview and do more harm than good.

What to avoid

  • Prompt screenshots. They show you typed something, not that the result was good.

  • Certificate collections. A free foundation badge from the government’s AI Skills Hub is worth a line on a CV, but it is a baseline, not leverage.

  • Confidential material. Putting a client’s documents into a personal AI account to make a portfolio piece is a confidentiality breach and a data protection problem. It also tells an employer something about your judgement. If in doubt, ask permission or rebuild the example with synthetic data. Check your contract and your employer’s policy for anything specific.

  • Hidden AI use. If you used AI to produce a work sample for an application, say so and explain your role. With 61% of large employers already noticing undisclosed use, concealment is a bigger risk than disclosure.

Presenting it

Keep the portfolio short. A two-page PDF or a simple web page with three case studies is enough, each linked to a fuller write-up if asked. On your CV, use one line per case with the measured result. In interview, be ready to walk through one case live, including the failure you caught. That conversation is usually where the evidence lands.

Bottom line

Show your work, not the tool. A real task, a baseline, a checking method, a failure you caught and a measured result add up to the most convincing proof of AI leverage you can offer. They are also the hardest for another candidate to fake.

Sources

From AI and Jobs: UK, October 2026