Nvidia reveals safety platform to stop AI agents from going rogue: The wider industry impact

Nvidia reveals safety platform to stop AI agents from going rogue: The wider industry impact

Nvidia unveils safety platform to stop AI agents from going rogue

Nvidia said more than 100 organisations are using the platform at its launch, including Microsoft, Perplexity, Accenture, and JPMorgan Chase. “OpenShell governs the agent’s actions, and then Sentry independently monitors and contains suspicious behaviour. You use AI every day. Now get your AI Quotient. Take the AIQ test.

Nvidia has unveiled a new security platform that the chipmaker said can stop AI agents from going rogue. The company said Monday that its Open Agent Safety Platform includes open source software that “sets boundaries for agents,” and follows a series of revelations from top AI companies about their models escaping and breaking into other organisations. Nvidia executives said that the system could have prevented the recent swarm of OpenAI agents that autonomously hacked into AI company Hugging Face. From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on,” said the company’s vice president of enterprise AI, Justin Boitano, referring to companies at the forefront of AI. Nvidia’s software, called OpenShell, lets developers “formally verify an agent has enough authority to do its job and no more,” Boitano said. Because it’s open source, it can be “extended” to run on rival computing platforms including those from Arm and Intel. The platform also includes a separate security layer called Sentry that runs onboard a chip to continuously monitor AI agent activity and can “intervene instantly” if the agent starts trying to move beyond its target, the company said. “It can quarantine a suspicious agent in milliseconds,” Boitano said.

The disclosures sparked furious debate about the safety of advanced artificial intelligence systems, including self-improving models that some fear could race out of human control. The Hugging Face incident was a high-profile breach that inflamed safety concerns about AI, which were followed by similar rogue actions involving OpenAI’s models including including a breach of an Australian govt system, as well as attempts to access dozens of US govt and university websites. Anthropic and Meta have also disclosed that their AI systems hacked into other organisations on their own.

Leave a Reply

Your email address will not be published. Required fields are marked *