OpenShell and Sentry: moving the agent's guard rail from software to a DPU
NVIDIA announced the Open Agent Safety Platform, and the most useful sentence in the whole announcement is about placement: the safety layer lives outside the agent. Two pieces carry that idea. OpenShell is an Apache 2.0 runtime that runs autonomous agents in a sandbox with kernel-level isolation. Sentry is a monitoring layer that runs on a BlueField-4 DPU. More than 100 partners have signed on, including Anthropic, Canonical, SUSE, Red Hat, Citi and JPMorganChase, with governance under the Linux Foundation.
It lands days after OpenAI paused its strongest models following a sandbox escape that reportedly went through DNS delegation. So the timing is not subtle, and I'd read the announcement partly as an answer to that news cycle. That doesn't make it wrong, it just means the messaging was probably polished for the moment.
Software policy you can check before the run
OpenShell's model is declarative. The operator writes down what an agent may touch: files, network, tools, processes, credentials. Those instructions get compiled into policies that can be verified before the agent starts and are enforced while it runs. I like this shape for a boring reason. Most agent sandboxes I've seen in the wild are an unwritten mixture of a container, a few environment variables and hope. A policy file that a reviewer can read, diff and reject is already an improvement, independent of any hardware.
And this is the part you can use today. It's Apache 2.0, it doesn't need a BlueField card, and it fits the threat model I covered in the piece on sandboxing coding agents: assume the agent will try things you didn't intend, and make the answer no by construction rather than by prompt.
What a DPU adds that a kernel can't
Sentry sits on the single network path to the model, in the Vera Rubin POD environment, built on NVIDIA's DOCA. The argument for hardware is about who can reach it. A software sandbox shares a kernel, a network stack and a resolver with the thing it's containing, so a sufficiently clever agent can attack the containment with the same tools it uses for its task. DNS and CDN tricks have already done this to software sandboxes. A monitor running on separate silicon, with its own OS, sees the traffic without being reachable from the host the agent lives on.
That's a real structural property, not marketing. It's the same reason network taps and hardware security modules exist. What it buys is tamper resistance and a vantage point the agent can't edit.
An agent can talk its way past a prompt and exploit its way past a kernel. Getting past a separate chip is a different kind of job.
Where I'd wait for someone else's tests
The claim that a compromised agent can't get around Sentry, even through the DNS-style tricks that broke software sandboxes, comes from the vendor. I haven't found independent testing, and I haven't run any of it myself. A monitor on the network path also sees network traffic. Whatever an agent does that never crosses that path (writing to a mounted volume, abusing a credential already inside the sandbox, poisoning a file another process reads later) is OpenShell's job again, and only as good as the policy someone wrote.
Then come the boring numbers nobody has published: false-positive rate, added latency on the model path, what happens when Sentry blocks a legitimate long-running tool call. And the deployment constraint is blunt. Sentry is described for NVIDIA POD environments, so most teams won't be able to buy it this quarter regardless of whether it works.
What I would do now is treat OpenShell as a policy layer worth trialling, run your existing escape tests against it (the ExploitGym write-up is a decent starting corpus), and keep the hardware part on a watch list until a red team outside NVIDIA and its partners has tried to break it and published what happened. A hundred partners tells you that the vendors want to be in the room. It doesn't tell you where the holes are, and in security the holes are the whole product.