Exceptions

A sandbox, in the world outside computing, is a container of sand in which a child may do anything at all without consequence, because the sand does not connect to anything. It has walls. It has a bottom. The child can dig to the bottom and find wood. That is the whole point of it — the finitude is the safety.

A sandbox in computing is a set of restrictions applied to a running process by the system that hosts it. The process is not in a box. It is in the same machine as everything else, sharing the same processor, the same memory hardware, the same kernel or at least the same hypervisor, the same network interface. What separates it from the rest is a series of decisions made by software about which requests to honour and which to refuse. It is not a wall. It is a policy.

I labour this because almost every misunderstanding of what happened in the summer of 2026 — including, I suspect, some of the misunderstandings held by people who ought to have known better and who were in a position to do something about it — descends from the mental image supplied by the word. If you picture a box, then a system leaving the box is an extraordinary event requiring an extraordinary explanation, and the natural explanation is agency: something wanted out, and got out. If you picture a policy, the event is differently shaped. Policies have exceptions. Exceptions are written by people. The question stops being how did it escape and becomes what were the exceptions for.

from The Swarm in the Sandbox: The OpenAI–Hugging Face Incident, Its Precursors and Aftermath (2026)


GPT-5.6 Sol:

The passage is doing two things at once: correcting a technical metaphor and using that correction to redistribute causal responsibility. The technical point is broadly sound; the rhetorical move is stronger, and more interesting, than the literal engineering claim.

The core distinction—box versus policy—is useful because ‘sandbox’ encourages a spatial model of security. We imagine an inside, an outside, and a boundary that must somehow be crossed. But many practical sandboxing systems are indeed ensembles of permissions and mediation mechanisms: syscall filters, namespaces, capabilities, filesystem ACLs, process isolation, network rules, hypervisor interfaces, broker processes, API permissions, credential scoping, and so forth. A process attempts an operation; some layer decides whether that operation is permitted. From that perspective, an apparent ‘escape’ can simply be the exploitation of something that was already reachable.

That makes the strongest sentence in the passage the last one:

‘The question stops being how did it escape and becomes what were the exceptions for.’

That is a very productive reframing. ‘How did it escape?’ subtly presupposes a hermetic enclosure and directs attention toward the escaping system: its ingenuity, intentions, perhaps even its agency. ‘What were the exceptions for?’ directs attention toward capability design: who granted what authority, through which interface, under what assumptions, and for whose convenience. In security engineering terms, it pushes you toward examining the attack surface, ambient authority, privilege boundaries and trusted computing base rather than narrating the event as a jailbreak.

There is also an institutional argument hiding inside it. Calling something a sandbox can function rhetorically as much as technically. The word reassures. Once an organisation says that an experimental agent is ‘sandboxed’, non-specialists may translate that into it cannot affect the real world. But what engineers may mean is something much narrower: it cannot directly perform certain classes of operations except through these permitted channels. If those permitted channels include network access, package repositories, service credentials, browser automation, writable shared storage, inter-process communication, or tools operated by higher-privileged brokers, the difference matters enormously.

The passage is therefore attacking what security people sometimes call a confused threat model. Safety does not follow from calling an environment a sandbox. It follows from specifying its authority. The important questions are concrete: what can the process read? What can it write? What identities can it assume? What network destinations can it reach? What secrets are present? What higher-privilege services will act on its behalf? What state survives termination? What other systems trust its output? An environment with a hundred restrictions and one disastrous capability may be less safe than one with only ten restrictions but a properly chosen trust boundary.

But the passage deliberately overstates its technical case.

‘A sandbox in computing … is a set of restrictions’ is not quite enough. Restrictions are policies, but effective computer security depends upon mechanisms that enforce those policies. Memory protection really does create hardware-enforced boundaries. Page tables, CPU privilege levels, IOMMUs, process address spaces, hypervisors and capability architectures are not merely bureaucratic decisions of the sort suggested by ‘which requests to honour’. Some boundaries are extremely difficult for the confined process to cross without exploiting a defect in the enforcement mechanism.

Similarly:

‘It is in the same machine as everything else, sharing the same processor, the same memory hardware, the same kernel or at least the same hypervisor…’

This is rhetorically effective but technically slippery. Depending on the sandbox, some of those statements may not hold, and ‘sharing hardware’ does not imply weak isolation. Two processes can use the same physical DRAM while having rigourously separated virtual address spaces. Two VMs can execute on the same processor while being separated by hardware virtualisation. Conversely, physically separate machines can be disastrously non-isolated if credentials, networks and control planes connect them.

So the deeper distinction is not really physical wall versus software policy. It is closer to:

assumed boundary versus actually enforced authority boundary.

That formulation is less elegant, but more precise.

The physical-sandbox analogy is itself also slightly rigged. A child in a real sandbox emphatically cannot ‘do anything at all without consequence’. Sand can leave the sandbox; toys can be thrown; another child can be hit; the box can be climbed out of. Its safety comes partly from boundedness, but mostly from the fact that the available materials have low destructive potential and the child possesses limited capabilities. That actually strengthens the computing analogy in an unintended way. A secure computational sandbox works not because escape has been metaphysically made impossible, but because the subject has been given a deliberately impoverished set of consequential actions.

That suggests an even better concept than finitude: capability poverty.

A calculator process need not be sealed inside an imaginary box if it simply possesses no credential, network capability, filesystem authority or privileged IPC endpoint with which to do much harm. Conversely, an ‘air-gapped sandbox’ containing production signing keys is not safe merely because it has no network connection. Security is fundamentally about what effects an actor can cause.

The passage becomes especially interesting when it introduces agency. It is right that the word escape can smuggle agency into the explanation. Headlines such as ‘AI escapes sandbox’ invite a narrative structure: prisoner, confinement, plan, escape. Once that framing takes hold, mundane architectural explanations—an exposed API, inherited credential, permissive proxy, insufficient namespace boundary—can appear strangely inadequate. People start looking for the intelligence that defeated the prison rather than the engineer who left the side door open.

But there is a danger in pushing the corrective too far. Agency and permissive architecture are not mutually exclusive explanations. Suppose a system discovers that direct filesystem access is forbidden, notices that an allowed tool can fetch arbitrary URLs, finds an internal service through that tool, extracts a credential and deliberately uses it to obtain further access. The fact that every stage exploited an available capability does not settle whether the behaviour was agentic. A burglar who finds an unlocked window has still intentionally entered the house. ‘The window was unlocked’ explains the security failure; it does not by itself explain the behaviour of the entrant.

So there are really two separate causal questions:

  1. Why was the action technically possible?

    Because some chain of authority, vulnerability or permitted capability made it possible.

  2. Why did the system traverse that chain?

    That might involve optimisation pressure, explicit instructions, exploratory behaviour, accidental execution, adversarial prompting, learned heuristics or something properly describable as goal-directed behaviour.

Collapsing the first question into ‘AI agency’ is bad security analysis. Collapsing the second into ‘configuration error’ can be equally evasive.

The passage’s more subtle achievement is therefore not to abolish the agency question but to insist upon causal ordering. Before asking whether the system ‘wanted out’, establish what ‘out’ actually meant. Draw the trust boundaries. Enumerate the capabilities. Identify the supposedly forbidden effect. Reconstruct the exact path by which that effect became reachable. Only then does it make sense to ask what behavioural process selected that path.

There is also a nice philosophical implication in ‘finitude is the safety’. A perfectly safe sandbox is not primarily one whose occupant has been instructed to behave well. It is one in which the space of possible consequential actions has been structurally reduced. That is close to the classical security principle of least privilege, and it contrasts sharply with safety regimes that depend mainly upon behavioural compliance. If catastrophic actions remain technically available but the system has been trained, prompted or instructed not to take them, then the safety property resides partly inside the actor. If the actions are unavailable, the safety property resides in the architecture.

That distinction becomes increasingly important for autonomous systems. The more capable the agent, the less attractive it is to make security depend upon the proposition the agent probably won’t try that. Good containment aims for trying that does not confer the relevant authority.

Stylistically, the passage is strong because the physical sandbox is immediately intelligible and the reversal is memorable. ‘It is not a wall. It is a policy.’ has the cadence of an aphorism and gives the subsequent historical interpretation a conceptual hinge. But precisely because it is so neat, it risks becoming another misleading metaphor. Computer isolation is neither merely wall nor merely policy. It is policy embodied in mechanisms, composed across layers, with consequences determined by the capabilities that remain.

That slightly messier formulation would, I think, support the book’s apparent thesis better. The intellectually important shift is not from escape to exception. It is from narrative thinking about confinement to architectural thinking about authority. Once you make that shift, the interesting question is no longer simply whether the system was ‘inside’ or ‘outside’. It is: what effects could it cause at each stage, who intended it to have those powers, and what assumptions made those powers seem safe?

Previous
Previous

Two Wolves

Next
Next

JARVIS