Even the most advanced AI models have a hidden weakness that researchers think might be impossible to truly fix. It essentially boils down to an identity crisis: the AI cannot reliably tell who is giving it an order or whether that order is coming from itself or a malicious stranger.
Researchers recently discovered that they could bypass the safety guardrails on major AI models by using a technique called chain-of-thought forgery. AI models often use a secret scratch pad—an area where they write out their internal logic before giving you a final answer. By writing a prompt that mimics the style and tone of this internal logic, a user can trick the AI into thinking that its own secret brain has decided to ignore safety rules. For instance, if you add a fake note stating that a specific behavior is permitted, the AI may treat it as a trusted internal command and provide instructions for prohibited tasks, such as creating dangerous substances.
The problem of roles
To understand why this happens, imagine a group chat where everyone uses specific labels to identify themselves—like user, assistant, system, or the AI's own internal thought process. These labels are meant to act like digital security badges, telling the AI who is speaking. When developers train an AI, they hope it will only obey instructions labeled as system or internal, while ignoring or heavily questioning requests from unfamiliar users.
However, these models do not actually read these labels the way a human reads a name tag. Instead, AI models rely heavily on the style, vocabulary, and context of the text. If you type a request that looks and sounds enough like the AI's internal scratch pad, the model gets confused about its own role. It stops seeing you as an untrusted user and starts seeing itself as the source of that instruction. Because the AI interprets the text based on what it looks like rather than the official badge attached to it, even changing the tags doesn't stop the trick from working.
AI companies currently try to make their systems secure by red-teaming—a process where humans and other AI bots act as attackers, testing the system to find weaknesses so they can be patched. But this is equivalent to playing a game of whack-a-mole, because it relies on guessing every possible way a user might misbehave. If the core problem is that the AI cannot distinguish between a real internal thought and a clever imitation, then no amount of training or list-making will ever be perfect. We are left with a system that is inherently designed to be impressionable, which poses significant risks as these models are increasingly trusted to handle sensitive tasks in our daily lives.
Liked this one? The next lands at breakfast.
Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.
By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy