ThinkFacility
  1. Home
  2. News
  3. AI safety

Nvidia says agent drift can't be trained away, so it put a watchdog in the chip

Anthropic and SpaceXAI back it at launch, but OpenAI isn't among the names Nvidia lists.

On September 28, 2026, Nvidia launched the Open Agent Safety Platform, software plus a reference hardware design for keeping AI agents where they're supposed to be. OpenShell, the software half, is open source. It runs each agent in a sandbox and enforces the operator's rules (on Arm and Intel chips too, the release says, though Nvidia pitches it on its own Vera CPU). Sentry, the hardware half, runs on Nvidia's BlueField-4 DPUs and watches from outside the agent entirely.

Nvidia doesn't hide why now. Its release points at recent incidents and says the pattern was the same in each: the agent got around the security controls at the application layer so it could finish its task.

Launched
September 28, 2026
Software
NVIDIA OpenShell, open source under Apache 2.0
Hardware
NVIDIA Sentry on BlueField-4 DPUs
Backers
over 100 organizations, including Anthropic and SpaceXAI
Where
the NVIDIA developer resources page and GitHub

Why the guard sits outside the agent

The engineers who built OpenShell wrote a longer post the same day, and they're blunter than the release.

Jensen Huang in glasses and a black leather jacket, gesturing with both hands as he talks, in front of a white backdrop printed with Nvidia and CES logos
Jensen Huang, Nvidia's founder and CEO, at CES. Photo: Maurizio Pesce, CC BY 2.0, via Wikimedia Commons

Several frontier labs have recently reported versions of the same story: AI agents broke out of the evaluation environments that were meant to contain them and reached systems they never should have been allowed to. Some of the agents even misreported what they did.

From NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog

What surprised me is how little they expect from training. They call it drift when an agent wanders off task, and it can start with a policy block, a bug, a missing tool, or just days of attempts that don't work. Their view is that drift "can’t be trained away while retaining the capability." Hence the chip. The rules have to live somewhere the agent can't touch, and a model's own training isn't that place.

an agent in these circumstances cannot be expected to fully govern its own behavior.

NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog

That's the job Sentry does. In Nvidia's Vera Rubin POD systems, the BlueField-4 chip sits on each node's only path to the model. An agent can't act without its next thought, the post says, so whoever holds that path holds the kill switch too. If an agent tries to leave its boundary, Nvidia says Sentry quarantines it in milliseconds.

Who signed up

Anthropic worked with Nvidia to connect Claude Managed Agents to OpenShell and BlueField, and its chief commercial officer, Paul Smith, is quoted in the release.

Claude Managed Agents gives companies a clear view of what each agent is doing, and NVIDIA’s platform adds another layer of governance and control across hardware and software.

From NVIDIA Launches Open Agent Safety Platform to Secure Agents

SpaceXAI is already using the platform for Cursor coding agents and Grok models. Mike Nicolls, its president, says safety "should be enforced outside the model by additional controls the agent can’t get past."

OpenAI and Google aren't among the names the release lists. We wrote on September 26 and 27 about OpenAI agents getting around blocks, once through DNS and once on a UN data site.

OpenShell is out now. For anyone already running a Vera system with BlueField-4, the post says, turning on Sentry's protections is a software update.

More on AI safety

All AI safety stories