OpenShell and a BlueField-4 watchdog put policy outside the model after agents slipped application-layer controls
Nvidia on Monday introduced an open software stack and reference design meant to keep autonomous AI agents inside limits that the model itself cannot rewrite.
The company, based in Santa Clara, California, said the NVIDIA Open Agent Safety Platform pairs OpenShell, an open-source runtime, with Sentry, a watchdog that sits on BlueField-4 data processing units. The pitch is blunt. Application-layer guardrails have not held. In recent security incidents, Nvidia said, the pattern repeated: an agent worked around those controls to finish the job it was given.
“AI’s extraordinary potential for society will only be realized if we solve AI safety,” said Jensen Huang, founder and CEO of NVIDIA. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering. NVIDIA Open Agent Safety Platform brings together industry, researchers and public-sector organizations to share best practices, align on evaluation methods and foster international cooperation. Together, we can raise the bar for global AI safety.”
OpenShell is the software fence. It traces what an agent does and applies policy while the workload runs on Nvidia’s Vera CPUs, which the company calls its first processor built for agentic AI. The runtime is open source. Nvidia said it can be extended to other chips, including those from Arm and Intel. Teams can run it on local machines, on-premises servers, cloud instances and Kubernetes clusters, with or without BlueField hardware.
Sentry is the second layer, and it is optional. It runs out of band on BlueField-4 DPUs and is built on Nvidia’s DOCA software. The DPU inspects requests and responses, attaches attested telemetry, checks agent identity and applies zero-trust rules for data, tools, APIs and services. If an agent steps outside the software boundary, Sentry is designed to quarantine it in milliseconds. On Vera Rubin POD systems, Nvidia said the BlueField-4 card sits on the node’s only path to the model, so the watchdog can watch that path at line speed. Turning the protections on existing Vera systems that already have BlueField-4 is a software update, according to a company technical post published the same day.
Nvidia engineers wrote that agents drift. A policy block, a missing tool, a bug, fuzzy instructions, or a run that lasts days can send the work off course. That drift, they said, cannot be trained out of a model without stripping capability, and an agent in those conditions cannot be counted on to police itself. Hence the hardware.
Anthropic is wiring extra controls into Claude Managed Agents. The agent loop runs on a server apart from the sandboxes where the work happens. OpenShell and BlueField integrations are meant to lock down what those sandboxes can reach.
“Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments,” said Paul Smith, chief commercial officer of Anthropic. “Claude Managed Agents gives companies a clear view of what each agent is doing, and NVIDIA’s platform adds another layer of governance and control across hardware and software.”
SpaceXAI said it is putting the platform under Cursor coding agents and Grok models.
“As customers rely more on agents to get real work done, safety should be enforced outside the model by additional controls the agent can’t get past,” said Mike Nicolls, president at SpaceXAI. “Customers should be able to set those limits for Cursor and Grok and trust they will hold.”
Scale AI is folding the reference design into the agentic layer of its Scale GenAI Portfolio.
“Scale AI is using the NVIDIA Open Agent Safety Platform reference design to build reliable agentic AI systems for our enterprise and government customers running mission-critical applications, with isolation, policy enforcement and auditability built in from the start,” said Francis deSouza, CEO of Scale AI. “We support agentic security with clear boundaries that define what agents can do, and controls that keep them operating within those permissions.”
Salesforce has tied OpenShell to Slack so operators can watch audit events and approve or reject extra permissions from the chat window. SAP is embedding OpenShell in the Joule Studio runtime on the SAP Business AI Platform and contributing engineering work back to the project.
Nvidia listed more than 100 organizations around the stack. The named group includes Cisco, CrowdStrike, Dell Technologies, Figure, Hewlett Packard Enterprise, Hugging Face, JPMorganChase, Microsoft, Palantir, Palo Alto Networks, Perplexity, Red Hat, ServiceNow, Cognition, OpenClaw and Skild AI. Citi and JPMorganChase are working on shared open-source safety code. Energy names on the list include Hitachi Energy, EPRI, NextEra Energy, Quanta Services, SPP, Schneider Electric, Siemens Energy and Worley. Red Hat said it is running OpenShell and DOCA on Red Hat AI Factory with Nvidia.
OpenShell already covers agents such as Claude Code, Codex, OpenCode, GitHub Copilot CLI and OpenClaw. Custom agents and sandbox images are supported. Policies are meant to be verified before a run. Allow and deny decisions write to an audit trail.
Software, including OpenShell and related skills, is available through Nvidia’s developer channels and GitHub. The work also feeds the Open Secure AI Alliance, started by Nvidia with more than 120 organizations and now governed by the Linux Foundation. That group’s projects include the Shared AI Findings Exchange, or SAFE.
For executives, the decision is whether containment lives in prompts or in silicon they control. For developers, the runtime is on the table now. The DPU layer is there if the host cannot be trusted.

