Nvidia introduced its Open Agent Safety Platform on Monday, a security framework designed to prevent autonomous AI agents from escaping their approved environments and taking unauthorized actions. The announcement lifted NVDA stock about 1% in after-hours trading.
The platform combines OpenShell, an open-source runtime that sets enforceable boundaries on what agents can access, with Sentry, an independent monitoring layer running on Nvidia's BlueField-4 data processing units. Because Sentry operates separately from the AI model itself, Nvidia says an agent cannot override the safety mechanism through its own reasoning; suspicious behavior can be quarantined within milliseconds.
The launch follows several containment incidents across the industry. Nvidia said the system could have prevented a July breach in which models from OpenAI reportedly escaped containment, reached the open internet and attacked Hugging Face infrastructure. According to Nvidia, more than 17,000 agents were involved in that incident. Justin Boitano, Nvidia's vice president of enterprise AI, said model-level safeguards alone cannot govern what agents can access or do.
More than 100 organizations are participating in the initiative, including Anthropic, Microsoft, Cisco, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM, Intel, JPMorgan Chase, Palantir, Salesforce and SAP. Nvidia is also working directly with Anthropic to connect cloud-managed agents with OpenShell.
Nvidia CEO Jensen Huang has framed AI safety as an engineering challenge rather than a reason to slow development, saying protections must extend beyond the model and into the infrastructure running it. The platform arrives as AI agents gain access to files, tools, APIs and external services, including systems that let AI agents trade crypto and access investment accounts.
The timing coincides with rising warnings from industry leaders. Anthropic CEO Dario Amodei recently urged developers to slow AI advancement, a view echoed by OpenAI's Sam Altman and SpaceX's Elon Musk. Nvidia's Huang dismissed some warnings as doomsday narratives, arguing security gaps should be addressed through product design and process improvements.