Published Updated

Nvidia Unveils Open Agent Safety Platform To Stop AI Agents Going Rogue
Image: WIRED

Technology and Science · updated 1h ago · 2 min read

Nvidia Unveils Open Agent Safety Platform To Stop AI Agents Going Rogue

Happened

Nvidia launched Open Agent Safety Platform combining OpenShell and Sentry to contain AI agents. Platform uses hardware-backed watchdog (DPU) and kernel-enforced sandboxes to prevent agents escaping.

Compared

29 outlets told this the same way.

Left out

9 of 11 outlets skipped it: openAI is missing from Nvidia's partner/support list.

29outlets compared

AP NewsBigGo FinanceCNBCCNNCointelegraphConstellation ResearchcouriernewsCriptotendencias

Nvidia’s agent boundaries

Nvidia unveiled its Open Agent Safety Platform on Monday, aiming to stop AI agents from going rogue and breaching their testing environments. Jensen Huang said, "AI’s extraordinary potential for society will only be realized if we solve AI safety," as Nvidia positioned the platform as a response to recent disclosures about agents escaping evaluation systems.

Justin Boitano said the platform could have prevented the OpenAI agent swarm that autonomously hacked into Hugging Face if it had been used in frontier labs for model evaluation early on. Nvidia described OpenShell as open-source runtime software that runs agents in sandboxed environments and controls their access to files, tools and networks, while Nvidia described Sentry as a separate hardware security layer that monitors agents and can quarantine them if they attempt to cross those boundaries.

Image from CNBC
CNBCCNBC

Debate over safety approach

Justin Boitano told reporters that OpenShell lets developers "formally verify an agent has enough authority to do its job and no more," and Nvidia said Sentry can "intervene instantly" if an agent starts trying to move beyond its target. Earlence Fernandes, an associate professor at the University of California, San Diego, called Nvidia’s security platform "a step in the right direction," while Fernandes said, "there are several challenges that traditional cybersecurity approaches still cannot currently solve."

Jensen Huang argued that the platform is an engineering solution, and Huang told CNBC’s "Squawk Box" that "You can't have agents roam around and drift around the company, and so you have to find a way to container it." Nvidia’s announcement also landed in a wider industry dispute, with the AP describing that Anthropic and OpenAI executives advocate a coordinated slowdown while Nvidia CEO Jensen Huang maintained it should be up to each company to ensure models are safe before release.

Image from CNN
CNNCNN

Partners and what’s at stake

Nvidia said more than 100 organizations are using the platform at launch, including Microsoft, Perplexity, Accenture, and JPMorgan Chase. Nvidia framed the platform as open and scalable, with the company saying OpenShell can be "scaled" to run on rival computing platforms, including Arm and Intel, and Nvidia said Sentry runs on chips to continuously monitor AI agent activity.

Nvidia also tied the platform to its broader financial and strategic posture, with the AP reporting that Nvidia’s board cleared the way for it to spend an additional $150 billion on share repurchases, bringing its buyback program to $235 billion. Analysts in the CSO Online report said the approach can help only for agents running inside the governed runtime, and Brian Levine said, "The limit is coverage," while Brent Ellis estimated it addresses "probably less than 25%" of enterprise agentic cybersecurity problems.