Nadella: treat AI models like an insider threat and build an "emergency brake"
Microsoft CEO Satya Nadella published a long post on X about controlling advanced AI models. He proposes separating the model from the system that controls it and assuming from the start that the model is compromised, meaning it may act against us, so it has to be surrounded with deterministic safeguards and procedures. An authorized person should always be able to stop the model mid-task, "like an emergency brake." Every significant action by the model should leave a tamper-resistant, human-readable trail, and transparency of reasoning is, in his view, "non-negotiable." "The most trustworthy superintelligence system will not be the one in which we trust the model the most, but the one that lets us trust it the least," Nadella writes.