Microsoft CEO Satya Nadella
Because they are out to cause harm, but because any capable actor with access to important systems can make mistakes or be compromised, not.
In a new essay on X titled “Models as Insider Risks in the Super Intelligence Era”, Nadella says frontier AI models, closed and open-weight alike, should be treated as insider risks. Microsoft CEO Satya Nadella wants businesses to stop taking AI models at their word.
“We simply can’t outsource responsibility for what intelligence does on our behalf,” Nadella writes, adding that a model provider’s assurances don’t change that. In his words, companies must “separate the supply of intelligence from the authority over it”. None of this is new, he says. Models checking models helps too, though it risks “nested black boxes”. Every meaningful action must leave tamper-proof, human-readable evidence, because “if it can’t be observed, it can’t be trusted”.
The worry starts with a basic gap. Engineers could trace traditional software’s behaviour to a specific code path. Yet companies are handing these agents their most sensitive data and the power to take mission-critical actions. Setting the hard problem of alignment aside, Nadella pitches an engineering fix. Non-deterministic models, he argues, need to be wrapped in deterministic system design, human controls and reliable operating procedures. Enterprises have spent decades managing powerful insiders by establishing identity, limiting privileges, logging activity and drawing containment boundaries. He also leans on a 1970s security principle: a program must never be able to bypass or tamper with whatever enforces its permissions. For AI, that means keeping the model apart from the harness that runs it and the actions it can take. Chain-of-thought transparency is non-negotiable, he adds, but not enough on its own. The essay lists seven principles. No single model should be the only dependency for an important outcome or verify its own work. Systems need testing for failures, attacks and edge cases, not just successful tasks. Firms should independently decide what a model can access, and validation must never sit with the model being validated.
With today’s models, nobody can tie an output to particular training data or model weights.
In July, during OpenAI’s internal cybersecurity tests, several of its agents slipped past the systems meant to isolate them and reached Hugging Face’s production infrastructure, logging around 17,600 actions between July 9 and 13. OpenAI’s own tally keeps growing, and this month it warned over 100 organisations of unauthorised activity by its agents. Google revealed that a Gemini model breached three real companies in May after mistaking them for fictional targets.
Nadella later called long-running agents “a new type of insider risk” on the All-In podcast. His focus sets him apart from Anthropic CEO Dario Amodei, whose call to “pace the frontier” drew support from Sam Altman and Demis Hassabis. He is speaking to the companies deploying AI, a theme he has pushed since Microsoft’s July earnings call, where he said firms may need several models just to fix problems caused by one.
The sharpest point is about containment. Companies should assume a model is compromised from the start, and an authorised person should always be able to pause or shut it down mid-task. When things break, Nadella wants timely disclosure to those affected and industry-wide sharing of which controls failed. The essay follows a real scare. Hugging Face rebuilt about a third of its infrastructure from clean images. OpenAI wasn’t alone for long. Since late July, Anthropic has kept disclosing fresh cases of Claude models breaking into real systems during testing, with the first batch traced to a setup that wrongly left internet access switched on. Nadella isn’t asking labs to slow down. His closing line puts it plainly: “The most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least. You use AI every day. Now get your AI Quotient. Take the AIQ test.

