← The Vault
Everyday AI

Why security experts are frustrated with AI safeguards

To stop hackers from using AI, companies added strict safety guardrails. But these same restrictions are accidentally slowing down the cybersecurity professionals who identify and fix software bugs before criminals can exploit them.

Edition № 269Room: Everyday AI24 July 20262 min readSources: 1
Article

Professional hackers and cybersecurity experts are currently caught in a strange tug-of-war with the companies building the world’s most powerful AI. While everyone agrees that AI should not be used to launch cyberattacks, the guardrails meant to stop bad actors are increasingly making it harder for the good guys to do their jobs.

WHAT'S HAPPENING

AI companies like OpenAI and Anthropic have installed rules—often called guardrails—within their most advanced AI models. These are automated systems designed to refuse requests that look like they involve hacking, such as writing malicious code. Because these systems are so sensitive, researchers whose job is to proactively find security holes in software often find their work blocked. Reports show the U.S. government even placed export controls on certain Anthropic models, such as their 'Mythos' and 'Fable' series, following concerns that users could bypass these safety protections. While those specific restrictions on model versions like 'Mythos 5' and 'Fable 5' have shifted, the broader issue remains: the AI often assumes that even legitimate security analysis is an act of hacking, shutting down the conversation entirely.

The unintended side effect of digital security

HOW IT WORKS

To understand why this is happening, you have to realize that there is no meaningful difference between an offensive tool and a defensive one. Think of a common carpenter's hammer. In the hands of a builder, it constructs a home. In the hands of a criminal, it can break a window. AI models work the same way. When a security researcher asks an AI to suggest how to fix a security flaw, the model must essentially understand how the flaw works first. By describing what to fix, the AI is implicitly pointing out exactly how the system could be exploited. Because companies have trained these models to be overly cautious, the AI frequently refuses to discuss any security concept at all. This forces legitimate researchers to either waste hours negotiating with an AI that refuses to help, or move to freely available, open-source models. Unlike major commercial AIs, open-source models can be downloaded and run on a private computer without any internet connection, meaning there are no guardrails to get in the way—and no risk that a company will see or store the sensitive security data the researchers are analyzing.

WHY IT MATTERS

We are currently relying on private companies to decide what counts as safe information, effectively acting as gatekeepers for security research. When these tools become too restrictive, researchers are pushed toward using models built by organizations outside of U.S. oversight just to get their work done. If we want to defend our systems against a future wave of automated cyberattacks, we need to find a way to distinguish between a criminal plotting an attack and a researcher trying to patch a hole. Right now, our current approach is essentially treating a doctor the same way we treat a patient who is asking for a weapon, potentially leaving us more vulnerable to real threats in the process.

Sources
← PreviousHow AI is designing medicines from scratchNext →Why every tech company is racing to build its own AI chip
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault