Cyber Security

Why your computer is less safe because AI companies are trying to protect it

AI guardrails are designed to stop hackers, but they are currently hindering the legitimate security researchers who keep our digital world safe.
Why your computer is less safe because AI companies are trying to protect it

Public opinion treats AI guardrails like a digital bulletproof vest. The logic is straightforward: if a model refuses to write malicious code, hackers have one less weapon in their arsenal. Behind the corporate marketing and government oversight, these safety measures are currently blinding the very people who protect our devices from real-world attacks. By trying to lock the doors against criminals, AI labs are accidentally locking out the locksmiths.

Technologists and security researchers are sounding the alarm on a trend that prioritizes optics over actual safety. While companies like Anthropic and OpenAI market their models as powerful engines of innovation, they have also wrapped them in layers of restrictive code designed to prevent "misuse." For the average user, this means the software you rely on—from your banking app to your smartphone operating system—is becoming harder to defend.

The myth of the digital bulletproof vest

The fundamental problem with AI guardrails is that cybersecurity is a two-way street. To fix a hole in a wall, a mason first has to understand exactly how that wall can be broken. In the digital world, this is called offensive research. Legitimate security professionals spend their days acting like hackers to find vulnerabilities before the real criminals do.

Chris Anley, the chief scientist at security consulting giant NCC Group, uses a simple analogy for this. A hammer is a tool you need to build a house, but it is also a weapon. You cannot build a structure without one. If you ban hammers to prevent violence, you also ensure that no houses get built. When a researcher asks an AI to try and exploit a bug, they are trying to prove a vulnerability is real so a manufacturer will fix it. If the AI refuses to answer because the request looks "malicious," the defender is left without their most efficient tool.

This creates a paradox where the same mechanism meant to hinder a hacker also stops a defender. The two roles use the exact same techniques. By restricting access to these capabilities, AI companies are making the jobs of network defenders significantly harder. This results in software flaws staying open for longer periods, which gives actual criminals more time to find and use them.

Negotiating with a tireless but stubborn intern

For many researchers, using a modern AI model is like working with a tireless intern who is also incredibly paranoid. You might ask the intern to analyze a piece of code for errors. The intern sees the word "exploit" or "vulnerability" and immediately stops working, giving you a lecture on ethics instead of the data you requested.

Chris Thompson, the chief executive of RemoteThreat, notes that researchers now spend a massive amount of their day negotiating with the model. Instead of focusing on the core security problem, they are trying to figure out why a model provided a helpful answer yesterday but refuses to speak today. This over-sanitization of AI output leads to inconsistent results.

In some cases, the guardrails are so sensitive that they trigger at the mere mention of security-related topics. A researcher working for a major smartphone-component manufacturer reported that their AI tools are almost useless because the guardrails are too strict. If the model catches wind of any security-related activity, it simply shuts down. This creates a productivity tax on the very people who are supposed to be protecting the global supply chain.

The bureaucracy of vetting programs

AI labs are not blind to this issue, but their solution is a layer of bureaucracy. Anthropic has its Cyber Verification Program, and OpenAI offers Trusted Access for Cyber. These are vetted programs where researchers can apply for special access to models with fewer restrictions.

However, these programs are often opaque and exclusionary. Access is frequently limited to specific Western organizations or those that meet strict corporate criteria. Even when a researcher is approved, the boundaries can shift without notice. This gatekeeping is particularly evident in the case of Anthropic’s Mythos model. The company has marketed Mythos as a high-end tool so powerful that it requires government-level oversight.

In June, the U.S. government briefly applied export controls on Mythos and another model called Fable. This move followed reports that the models' guardrails could be bypassed to execute cyberattacks. While the controls were later eased for vetted U.S. organizations, the incident highlighted a growing trend of treating AI as a controlled weapon rather than a general-purpose tool. This level of gatekeeping does not stop hackers who use their own hardware, but it does create a massive hurdle for legitimate companies trying to stay ahead of the curve.

The migration to open source and foreign models

When legitimate researchers are blocked by U.S.-based AI companies, they do not simply stop working. Instead, they look for tools that do not have these restrictive filters. This often means moving away from polished cloud models toward open source alternatives that can run on private hardware.

Paolo Stagno, the chief technology officer at Crowdfense, explains that his team avoids cloud-based models for sensitive work. Feeding a secret software flaw into a cloud AI risks leaking that data or having it absorbed into a future training run. To maintain privacy and avoid the "babysitting" of corporate guardrails, many professionals are now using open source models that they can run locally without any oversight.

There is a strategic risk here. If U.S. researchers find that domestic models are too restrictive to be useful, they may turn to powerful open source models developed in other countries. Chris Thompson points out that researchers are already being pushed toward Chinese models like GLM. These models are freely downloadable and come with no usage restrictions. This shift moves the center of innovation away from U.S. governance and into environments where Western safety standards do not apply at all.

Why cloud models are a privacy nightmare for security pros

Privacy is a foundational concern for anyone working in high-stakes cybersecurity. If a researcher finds a "zero-day"—a flaw that the software manufacturer doesn't know about yet—that information is worth millions of dollars on the open market. It is also a critical piece of intelligence for governments.

Sending that information to a server owned by OpenAI or Anthropic is a massive risk. Even with specialized access programs, there is no guarantee that the data won't be used to train the next version of the AI. Some researchers, like Giuseppe Cali, use AI only for the boring parts of the job, such as translating complex code into a readable format. They refuse to use it for the actual discovery of bugs because they want to keep their work private.

This caution means that the most advanced AI models are currently being used for low-level tasks rather than the high-level reasoning that could actually improve global security. The guardrails and the cloud-based nature of these tools create a ceiling for how much they can actually help the people protecting our networks.

What this means for your digital safety

The bottom line for the average consumer is that our digital defenses are not keeping pace with the speed of AI-assisted attacks. Criminals do not follow the rules and do not use vetted programs. They use uncensored, open source models or their own custom-built AI to automate the process of finding and hitting targets.

Meanwhile, the "good guys" are stuck in a cycle of applying for permissions and arguing with chatbots. Practically speaking, this means the software on your phone and laptop may have vulnerabilities that stay unpatched for longer. It also means that the cost of cybersecurity is rising, as firms have to spend more time on manual work that AI could theoretically handle in seconds.

To keep the digital world safe, AI labs need to move away from the idea that they can babysit every user. Instead of building walls that block everyone, they need to provide responsible, high-speed access to the people who are actually on the front lines. Without that shift, we are essentially asking our security researchers to win a high-speed race while we are the ones tying their shoelaces together.

Key takeaways for the everyday user

  • AI is a dual-use tool: The same AI that helps a doctor diagnose a disease can help a hacker find a flaw in a bank's security system.
  • Guardrails are not perfect: While they stop some low-level abuse, they are easily bypassed by determined criminals using private hardware.
  • Defenders are being slowed down: The people who find and fix bugs in your apps are struggling to use AI because the safety rules are too broad.
  • Privacy is a hurdle: Professional researchers often avoid the most popular AI tools because they don't want to leak sensitive security data to the cloud.
  • Open source is the alternative: The industry is moving toward local, unrestricted models to get work done without corporate interference.

Sources:

  • Anthropic Public Disclosures on Mythos and Fable Model Access
  • OpenAI Trusted Access for Cyber Program Documentation
  • NCC Group Technical Analysis on AI in Offensive Security
  • TechCrunch Report on AI Guardrails and Cybersecurity Researchers
bg
bg
bg

See you on the other side.

Our end-to-end encrypted email and cloud storage solution provides the most powerful means of secure data exchange, ensuring the safety and privacy of your data.

/ Create a free account