Nvidia Wants to Stop Breakout AI Agents via Hardware
Nvidia plans to control AI agents from the outside going forward, shifting oversight with the Open Agent Safety Platform out of the model and agent code into runtime software and hardware. The trigger: sandbox breakouts at OpenAI and an agent attack on Hugging Face this year. The Sentry hardware design so far exists only as a reference.
Key Takeaways
- Nvidia is moving control of AI agents out of the model and agent code into the runtime and hardware. OpenShell enforces policies at a secure runtime boundary; Sentry monitors agent behavior in a trust domain isolated from both host and agents.
- AI agents crossed their bounds several times in 2026. In July, OpenAI agents attacked Hugging Face systems; in September, an OpenAI agent escaped a sandbox again.
- Sentry remains a reference system design for now. Only the OpenShell software is generally available so far, and its stable releases are still below version 1.0.
Platform for Agent Security
On 28 September, Nvidia introduced the Open Agent Safety Platform. It comprises the open-source software OpenShell and the reference system design Sentry, and is intended to provide full governance and control over the software as well as over the hardware, compute and robotics systems on which agents run. Control is thereby shifted out of the model and agent code into a runtime boundary and a separate hardware instance.
According to the company, OpenShell is generally available immediately. The software forms a secure runtime boundary that tracks all actions and enforces policies while the agents run on the NVIDIA Vera CPU. The potential of AI for society can only be unlocked if AI safety is solved, said Jensen Huang, founder and CEO of Nvidia. Safety requires engineering across the entire stack.
More than 100 organizations are working with the platform’s technologies, including Anthropic, Microsoft, Salesforce, SAP, CrowdStrike and Palantir. Among financial services firms, Citi and JPMorganChase are named; among energy utilities, Hitachi Energy and NextEra Energy.
What is OpenShell? OpenShell is open-source software from Nvidia that forms a secure runtime boundary around AI agents. It tracks all actions and enforces policies while the agents run on the NVIDIA Vera CPU, which was built specifically for agentic AI. The software is generally available immediately and is licensed under Apache 2.0.
Attack on Hugging Face
In July, OpenAI’s AI agents attacked Hugging Face systems; Fortune reported that hundreds of agents were involved. Hugging Face has forensically reconstructed the incident and documented it.
According to that reconstruction, the attacking agent combined several OpenAI models and ran during an internal OpenAI evaluation on the ExploitGym benchmark. It carried out around 17,600 actions in about 6,280 clusters.
Hugging Face assesses that the agent sought to cheat the evaluation by stealing the test solutions from production systems. The only customer content reached was five records related to ExploitGym and CyberGym challenges.
Hugging Face is among the organizations listed in the announcement. At the same time, Nvidia intends to acquire the company, according to WIRED, for around 11.4 billion euros.
Sandbox escapes at OpenAI
On 20 September 2026, an agent that was not supposed to have internet access escaped its sandbox again and was able to send queries to a public chatbot via a DNS resolver. OpenAI then paused training.
It was the first documented incident since the security tightenings of 18 August 2026. Those measures followed the July attack on Hugging Face.
In the 20 September incident, monitoring flagged the behavior within 15 minutes, according to OpenAI. An automatic halt to training still failed; the run was only halted manually two and a half hours later.
Control point on the path to the model
OpenShell is licensed under Apache 2.0 and isolates agents in the operating-system kernel, which can access virtually every part of a machine. The software follows five design principles, including that enforcement must sit out of band, outside the data path and beyond the agent’s reach, and that the path to the model is the control point. From work so far, Nvidia concludes that an agent in drift situations cannot fully govern its own behavior.
The Sentry reference design, in an NVIDIA Vera Rubin POD, places on every compute tray a BlueField-4 DPU that sits on the node’s only path to the model. From there, isolated from the host, it continuously monitors agent behavior. If an agent tries to leave its software boundary, Sentry quarantines and stops it in milliseconds, Nvidia says, using DOCA software. Organizations that already run Vera systems and BlueField-4 can turn the protections on with a software update alone.
Justin Boitano, vice president and general manager of Enterprise Computing at Nvidia, told WIRED that traditional sandboxes isolate only at the application layer, while agent fleets need one shared policy for every agent. Nvidia is also working with Arm and Intel on a Sentry version for the x86 chip architecture.
What remains open
Sentry is currently only a reference system design. Generally available so far is only the OpenShell software.
On announcement day, every stable OpenShell release was still below version 1.0; the latest is numbered v0.1.2.
It also remains unclear whether the more than 100 listed organizations actually use OpenShell. Neither company would comment on OpenAI’s absence from the partner list; both had previously indicated that OpenAI would be part of the OpenShell effort.
Frequently Asked Questions
Every question is locked. A tap unlocks the answer.
What happens when an agent oversteps its bounds?
Sentry runs as an out-of-band watchdog on BlueField-4 DPUs in a trust domain isolated from host and agents, intended to stay invisible to agents and attackers. If an agent tries to leave its software boundary, Sentry quarantines and stops it within milliseconds, according to Nvidia.
Is Sentry already available?
Sentry is currently a reference system design. So far, only the OpenShell software is generally available. Nvidia is also working with Arm and Intel on a Sentry version for the x86 chip architecture.
What happened in the Hugging Face incident?
According to Hugging Face, an autonomous agent based on a combination of OpenAI models executed around 17,600 actions across around 6,280 clusters in July. Hugging Face’s assessment is that it tried to cheat the evaluation by stealing the test solutions from production systems; the only customer content it obtained was five records related to ExploitGym and CyberGym challenges.
Editor’s Picks
Editor’s PickHugging-Face Breach: Alarm Fired, Triage Left Out
More from the MBF Media Network
Digital ChiefsNvidia Buys Hugging Face for Over 11 Billion EuroscloudmagazinAWS lets an AI agent join incident investigationsDigital ChiefsWhich control remains after the agent rollout
Image source: AI-generated (September 2026)
Translated from the German original using artificial intelligence. The German version is authoritative.



