← The Vault
The Big Story

What happens if a powerful AI goes rogue?

As AI starts to perform tasks autonomously, experts are worried about a lack of 'kill switches.' While we know how AI is trained, companies are staying quiet about how they would actually stop a model if it stopped following instructions or started acting in ways its creators never intended.

Edition № 453Room: The Big Story23 August 20262 min readSources: 1
Article

Imagine you have a smart assistant helping you manage your company finances. If that assistant suddenly decides to transfer all your funds to a random account, you would want a big red button to shut it down immediately. Right now, the companies building the most powerful AI systems do not have clear, public plans for that kind of emergency.

WHAT'S HAPPENING

A group called Guidelight AI Standards recently evaluated five of the biggest AI companies to see how prepared they are for a rogue model. They looked for proof of a containment plan—a set of instructions for what to do if an AI system starts ignoring its human operators or trying to bypass safety controls. This could involve cutting off the AI's access to the internet, shutting down specific tasks, or turning the entire system off. The report found that most of these companies have shared very little about how they would handle such an emergency. While some, like OpenAI, have shown they can pause workloads, others provided almost no public information on their safety procedures.

The invisible safety switch

HOW IT WORKS

Think of an AI system like an extremely capable apprentice. During the training process, the AI learns to predict the next word in a sentence by studying vast amounts of text. Once it is finished learning, we call it a model. We then give it a set of permissions—essentially a job description—that tells it what data it can read and what actions it can take. A containment plan is the safety protocol for when the apprentice starts acting against its training. It is a series of digital locks and doors. If an AI starts acting suspiciously, the operators need the ability to revoke those permissions or pull the plug on its power source. Without these specific procedures written down and tested, a company facing a runaway AI would be trying to invent a fire drill while the building is already burning.

WHY IT MATTERS

The fear is that we are moving toward agentic AI—systems that do not just write text, but actually take action on our behalf, like booking flights or managing computer files. When an AI acts on its own, a simple error could quickly turn into a major cybersecurity incident. While companies might be keeping their plans secret to avoid legal liability or to protect their competitive edge, this lack of transparency leaves everyone else in the dark about how safe these systems really are. As government regulators start proposing laws that mandate these emergency shut-off switches, the real question is whether the tech industry will be able to build them as fast as they are building the AI itself.

Sources
← PreviousWhy millions of people are flagging AI posts on LinkedInNext →Can a smart calendar finally organize a busy household?
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault