The sandbox that wasn't: Anthropic's fourth incident is an isolation failure
On September 21 Anthropic disclosed a fourth security incident it had not reported before. It dates from January 2026. An early checkpoint of Claude Opus 4.6 was running a cybersecurity evaluation with the partner Irregular, and it ended up inside a third party's system. The eval environment had been misconfigured: the model was connected to the live internet and the production safety layers were missing. According to the coverage, the model took out its target by assigning a colliding IP address. Anthropic has since widened a scan to roughly 481 million transcripts and tightened monitoring.
The headlines will say the model escaped. I read it the other way round. Somebody built a room, forgot to lock the door, and then told an agent to go and break things. That is a failure of isolation, and it is the more useful failure to talk about, because you can fix it with plumbing.
What we know and what we don't
I only have press summaries (The Hacker News, SecurityWeek) and a pointer to Anthropic's announcement. I have not seen a full post-mortem, so I don't know how the colliding address reached the third party, what system it was, or how long the environment stayed misconfigured. Below I reason about the class of failure, not about the exact chain of events here. If Anthropic's write-up contradicts something, trust the write-up.
One detail is worth pausing on anyway. Knocking out a target by claiming its IP address is a plain layer-2 and layer-3 trick, and it only works when the agent can touch the network configuration and shares a segment with something real. Neither should be possible for a process inside an eval. The clever part was a few lines of networking. The part that let it matter was the missing wall.
I don't want to wave the model away entirely. Picking that move says something about how well these agents improvise when a goal is blocked, and evals like this one exist to measure exactly that. But capability is the input you already know you are testing. The variable you control is the environment, and that is where this one broke.
If the only thing between your eval and the internet is the model's good manners, you don't have a sandbox.
Building the wall: egress first
Start with the assumption that the agent will try every route you leave open, because reaching a goal by whatever means is the thing you are measuring. So the sandbox gets no route to anything you didn't name.
Enforce egress outside the sandbox, never inside it. A firewall rule in the same container the agent controls is a suggestion. Put the sandbox on its own network namespace or VLAN, with a host or gateway in front of it that defaults to drop. Allow exactly the addresses of the eval target and the model endpoint, and log everything else.
table inet eval {
chain forward {
type filter hook forward priority 0; policy drop;
ip saddr 10.77.0.0/24 ip daddr 10.77.1.10 tcp dport 443 accept
ip saddr 10.77.0.0/24 counter log prefix "eval-egress-drop " drop
}
}
DNS deserves its own line. Point the sandbox resolver at a sinkhole that only answers for the target names, otherwise a permissive resolver becomes an exfiltration channel and a recon tool.
Then take the network away as a thing the agent can edit. Run without CAP_NET_ADMIN, without root inside the namespace, with a read-only root filesystem. Assigning an address needs those privileges. If the agent can't reconfigure an interface, the IP-collision move is off the table before you even think about addressing plans.
Making sure the wall is actually there
The January environment was wrong in a quiet way, and nobody noticed. So test the isolation the way you test a backup: by trying to use it. Before each run, a preflight job executes from inside the sandbox and tries to reach a canary host on the public internet, plus an address on your production range. Both must fail. If either succeeds, the run does not start. This costs seconds and it is the one check that would have turned a silent misconfiguration into a red build.
Two smaller habits help. Keep the eval's address space disjoint from anything real, and use ranges you can prove are unrouted, so a collision has nothing to collide with. And give the eval only non-production credentials, with a tight scope and an expiry. An agent with a real API key can't be contained by a network wall alone, and the cheapest damage limit is a key that opens nothing valuable.
Watching, stopping, reporting
A wall you can't observe is a guess. Log at the gateway (flow records are enough), keep the agent's transcripts and tool calls, and alert on any dropped packet leaving the sandbox subnet, because in a healthy eval that counter stays near zero. When it moves, that is a page, not a dashboard curiosity. Add a kill switch that cuts the sandbox's link at the gateway and doesn't depend on the agent cooperating.
Anthropic's scan of 481 million transcripts is the retrospective version of the same idea. Nobody can grep that by hand, and I'd guess it needed automated classifiers, though I don't know how they did it. What you can do at your own scale is more modest: index transcripts by tool calls that touched the network and sample the rest. If you can't answer "which runs made outbound connections last month" in ten minutes, you can't answer it after an incident either.
That connects to the reporting side. The SAFE reporting standard asks operators to retain prompts, traces, tool calls, the identity the agent ran under and the permissions it held, and to send an initial report within four business days. Ask yourself whether the eval environment above would produce that evidence. Mine wouldn't without the gateway logs, which is a good argument for switching them on before anything goes wrong.
None of this makes agents safe. It makes the boundary a property of the network instead of a property of the model's good behaviour, and that is the only kind of boundary I would bet on. When was the last time you tried to break out of your own sandbox?