HomeAIHow AI guardrails hinder the work of offensive cybersecurity researchers

How AI guardrails hinder the work of offensive cybersecurity researchers

How AI Guardrails Are Impeding the Work of Offensive Cybersecurity Researchers

For months, AI giants have been developing special, vetted programs and strict protections to limit the use of their models by malicious hackers. But these restrictions are now hindering the work of legitimate network defenders and offensive cybersecurity researchers.

In June, the US government imposed export control restrictions on Anthropic’s much-lauded Mythos and Fable AI models. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to create and carry out malicious cyberattacks.

Anthropic’s Marketing and Government Response

Regardless of whether the incident was truly motivated by fear of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cyber machine that can only be made available to carefully vetted users and even then with strict security measures in place. (Export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 was made generally available again on July 1; Mythos 5 was reintroduced only to vetted US organizations as part of the government’s review process.)

Gatekeeping in AI Models

This type of gatekeeping is not unique to Mythos. Both Anthropic, with its other models, and OpenAI offer programs for cybersecurity researchers to apply to be vetted and, if approved, access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber ​​program and Anthropic’s Cyber ​​Verification Program.

These guardrails have been widely criticized, particularly by researchers whose job it is to find unknown vulnerabilities in systems and find ways to exploit them before criminals do.

Researchers’ Perspectives on AI Guardrails

During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said: “It doesn’t really sit well with me that these randomly selected large companies are making arbitrary decisions about what is safe and what isn’t when it comes to security.”

Dowd has spent decades finding “zero days” — previously unknown software flaws and the exploits that exploit them — and selling them to Western governments rather than reporting them to software makers to be patched. Governments pay a premium for vulnerabilities precisely because they remain open, which is useful for intelligence operations.

Dowd admitted that his work may make him biased, but he is not alone. Several people who work in offensive cybersecurity — they proactively examine systems for vulnerabilities — described to TechCrunch how they use AI tools and deal with their guardrails.

Offensive vs. Defensive AI Tools

Chris Anley, senior scientist at security consulting giant NCC Group, said asking an AI model to exploit a flaw is an important step in confirming that it is a real vulnerability worth fixing. But if a guardrail causes the model to refuse to answer the question directly, the guardrail harms defenders, he said.

“This is where the whole offense versus defense and guardrails part comes into play, because ‘Fix this code’ as a prompt is both an essential defense mechanism and a guide to finding critical vulnerabilities in the code base,” Anley said. “So the same tool is both an offensive and a defensive tool at the same time, and the two can’t really be separated.”

It was “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool, but it’s also an unbreakable weapon.”

When he and his colleagues encounter such a hurdle, they sometimes turn to open-source AI models that have no guardrails.

Challenges with Guardrails and Open-source Alternatives

Paolo Stagno, the chief technology officer at Crowdfense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that AI companies with their vetted programs and guardrails are “essentially treating customers like children in need of babysitting.”

Stagno said he and his colleagues do use boundary models – but only for reverse engineering. They avoid using AI to find vulnerabilities or create exploits, he said, because feeding that work into a cloud-based model risks sensitive vulnerability data being lost or incorporated into future training runs. For this step, he said, they would use open-source models running locally because they do not rely on sharing data outside the model.

Giuseppe Cali, a security researcher who finds zero-day attacks and develops exploits, said guardrails don’t hinder his work. That’s because it doesn’t use AI for attack tasks; Instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. To that end, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.

“I still want to be in control of the actual detection and weaponization of bugs, and that wouldn’t change if all protections were lifted tomorrow,” Cali said. “I’m jealous of my bugs and I like this game too much to let models play it for me.”

A researcher at a smartphone component maker, who spoke on condition of anonymity because he is not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program and therefore its tools are of little use in finding vulnerabilities because the guidelines are too strict.

“If it gets windy, we do something safety-related, it just stops and can’t be used,” the person said.

Inconsistencies and the Push Toward Foreign Models

Chris Thompson – CEO of cybersecurity company RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI – said that in his experience with Frontier AI models, the guardrails can be inconsistent and work differently every day. This is true even within the looser boundaries of the tested programs from Anthropic and OpenAI.

“I think the practical impact is that you spend a lot of time negotiating with the model instead of working on the nuclear safety program,” Thompson said. “Instead of analyzing a vulnerability and justifying exploitability, try to figure out why you’re getting inconsistent results or why models are over-sanitizing the output.”

As a result, researchers are relying on or being pushed toward Chinese open-source models like GLM — freely downloadable models that can be run locally without review or usage restrictions — Thompson said.

“There are these responsible researchers who are being pushed out of U.S.-managed systems and into foreign-owned systems,” he said. “I think having these guardrails does more harm than good.”

Call for Responsible Access and Accountability

Instead of further tightening restrictions, Thompson called on AI Frontier Labs to open up their programs, provide responsible access, and hold accountable those who misuse their tools. Otherwise, he argued, defenders would lose the AI ​​race.

“There’s this big storm coming. There’s this big wave of attacks that are going to happen at an unprecedented speed and scale,” Thompson said. “But the same security consulting firms and legitimate researchers trying to make a difference are currently being suppressed.”

If you purchase through links in our articles, we may receive a small commission. Our editorial independence remains unaffected.

Source: Here

“`

Must Read
Related News

LEAVE A REPLY

Please enter your comment!
Please enter your name here