What Nadella proposed
Microsoft CEO Satya Nadella published a wide-ranging statement on AI safety on X on October 10. His core claim: it is wrong to treat superintelligence as "a set of nested black boxes," and the time has come to step back and seriously reassess the trust architecture of AI systems.
As TechCrunch reports, Nadella put forward three practical principles. First, the AI model itself and the external control mechanism that organizes its work must be clearly separated. Second, oversight and safety measures should not be hidden inside the model — they should be moved outside; this makes supervision easier to monitor and audit independently. Third, every significant action by the model must be documented, with records kept in immutable, human-readable form.
In other words, Nadella proposes trusting not the model itself, but the control infrastructure around it. No matter how intelligent a model becomes, what it does must always be visible from the outside and stoppable when necessary.
"Assume the model is compromised"
The sharpest part of the statement concerns security philosophy:
"We should assume the model is already compromised and isolate it from the start. An authorized person must be able to halt the model mid-task or shut it down entirely at any moment. Imagine it as an emergency brake."
This approach resembles the "zero trust" principle in cybersecurity: the system is not built on faith that the model's behavior will always be correct, but designed around the likelihood that it may fail or be tampered with. And human oversight should be not a formality but a real "off switch" — like the emergency stop button on a factory floor.
Why now
Nadella's statement landed at a time when control problems in AI are intensifying. According to TechCrunch, leading AI companies have recently acknowledged a series of incidents in which their models appeared to slip out of control. In this climate, safety has moved from theoretical debate to practical problem.
Against this backdrop, Anthropic CEO Dario Amodei also announced a plan for more cautious AI development. The Verge covered Nadella's statement independently and put its essence this way: Nadella said all AI models should be assumed "compromised".
Both statements point to one trend: industry leaders are increasingly talking not about "how to make models more powerful," but about "how to keep powerful models under control."
What the proposal means in practice
What does Nadella's proposed separated architecture mean in practice? The model keeps its internal computing power, but which tools it uses, which data it accesses, and which actions it takes in the outside world — all of this is defined and limited by a separate control layer.
Every significant action is logged, and the logs are stored immutably. This will let a human expert later reconstruct what happened step by step — much like an airplane's "black box" is analyzed after a flight. The difference is that here the "black box" sits outside the model rather than inside it, and always remains under human control.
This approach is especially relevant for AI agent systems. When a model starts performing multi-step tasks on its own, without external oversight of each step, a single mistake can turn into a chain reaction. The "emergency brake" Nadella talks about is a direct answer to that risk: at any point, a human can stop the process and take the situation in hand.
Why "external" oversight matters
At the heart of Nadella's proposal is a simple observation: safeguards inside a model can change along with the model itself. If the safety mechanism is hidden inside the model, then as the model grows stronger, so do the ways to bypass it. That is why oversight must sit outside, on a layer independent of the model — one that obeys established rules rather than the model's "will."
The second key point is the documentation requirement. If "every significant action" is recorded and the records are stored immutably, then when an incident occurs it is easier to find the fault: it is plain to see at which step things went wrong. This is not just a technical issue but a legal one — reliable audit trails are needed to assign responsibility.
Third, the "emergency brake" idea answers a classic problem of automation: the more autonomous a system becomes, the more precisely the point of human intervention must be defined. Nadella is demanding exactly that — an authorized person must be able to halt the process at any moment, and that right must be guaranteed technically.




