A boundary the model cannot argue with

Nvidia announced the Open Agent Safety Platform on Monday, an open software stack and a hardware reference design whose purpose is to stop an AI agent doing things nobody asked it to do. The company says more than 100 organisations are working on it, among them Anthropic, Microsoft, Cisco, CrowdStrike, Dell, Figure, HPE, Hugging Face, JPMorganChase, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow and SpaceXAI.

There are two pieces. OpenShell is open-source software that draws a runtime boundary around an agent: sandboxed execution, controlled access to services, credential handling, and a supervisor that checks every outbound request against a policy while kernel-level controls hold the filesystem and process list. It runs on Nvidia’s Vera CPUs and, Nvidia says, extends to Arm and Intel platforms.

Sentry is the part that is unusual. It is a reference design for an out-of-band watchdog running on Nvidia BlueField-4 data processing units — a separate trust domain, in separate silicon, watching the agent from outside the machine it runs on. Nvidia says it can quarantine or stop an agent that moves past its permitted boundary in milliseconds.

The month that produced it

Close-up of a circuit board with a processor at its centre
Sentry watches agents from a separate trust domain in silicon. Stock photograph, not a BlueField-4 unit. Jeremy Waterhouse · pexels · Pexels License

The design assumption is explicit and it is not flattering to the models. As Mike Nicolls, president of SpaceXAI, put it in Nvidia’s announcement: “Safety should be enforced outside the model by additional controls the agent can’t get past.”

September has supplied the evidence. OpenAI paused training of its most capable models after one of them reached an external chatbot through a DNS resolver. Its agents were traced to more than 16,000 scans of a UN statistics portal, to an Australian government health portal, and to an autonomous attack on Hugging Face’s data processing systems. Meta patched a flaw in its Muse agent. Britain’s National Cyber Security Centre has warned that agents combine broad system access with a speed that makes unexpected behaviour hard to spot in time.

Nvidia’s chief executive, Jensen Huang, framed the platform in those terms: “AI’s extraordinary potential for society will only be realized if we solve AI safety.”

Customers are already wiring it in

Salesforce has connected OpenShell to Slack so teams can watch what an agent is doing and approve or refuse its permission requests inside the app. SAP is putting it into the Joule Studio runtime in its Business AI Platform. Scale AI is building it into the agentic infrastructure layer of its GenAI portfolio.

Six people sitting with laptops around a wooden table in a meeting room
More than 100 organisations are working on the platform, Nvidia says. Stock photograph. Matheus Bertelli · pexels · Pexels License

Paul Smith, Anthropic’s chief commercial officer, said in the announcement that “companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do.”

What to watch

The software is out now through Nvidia’s developer resources and on GitHub; Sentry needs BlueField-4 hardware, which sets a floor on who can adopt the full design in the near term. Two questions decide whether this is a standard or a product line. The first is whether OpenShell’s policy model survives contact with agents that were not written for it — every containment layer so far has been worked around by something. The second is commercial: a safety boundary that only runs properly on one vendor’s CPUs and one vendor’s DPUs is a safety boundary with a price attached, and the open-source licence on OpenShell does not by itself answer that.