source: techcrunch ai: how ai guardrails are impeding the work of offensive cybersecurity researchers

level: technical

ai companies like anthropic and openai have set up vetted programs and guardrails to limit how their models can be used for cyberattacks. these restrictions are now slowing down offensive cybersecurity researchers who find and exploit vulnerabilities before criminals do. mark dowd, a well-known security researcher, said it is uncomfortable that large companies decide what is safe in security. chris anley, chief scientist at ncc group, explained that asking a model to exploit a bug is a key step in confirming a real vulnerability. when guardrails block that, it hurts defenders because the same tool is both offensive and defensive.

some researchers bypass the restrictions by using open-source models that have no guardrails. paolo stagno, cto at crowdfense, said ai companies treat customers like children with their vetted programs. his team uses frontier models only for reverse engineering, not for finding vulnerabilities, to avoid leaking sensitive data to cloud-based models. giuseppe cali, a security researcher, said guardrails do not impede his work because he uses ai only for initial reverse engineering and tool building, keeping bug discovery for himself. an anonymous researcher at a smartphone-component maker said anthropic's strict guardrails make its tools barely usable for finding vulnerabilities.

chris thompson, ceo of remotethreat, said guardrails are inconsistent and force researchers to spend time negotiating with the model instead of analyzing vulnerabilities. this pushes some toward chinese open-source models like glm, which can be run locally without restrictions. thompson argued that responsible researchers are being pushed away from u.s.-governed systems, which is more harmful than good. he called for ai labs to open up programs, provide responsible access, and hold abusers accountable, warning that defenders will lose the ai race if stifled.

why it matters: overly strict ai guardrails can slow down legitimate security research, pushing experts toward unregulated foreign models and potentially weakening cyber defenses.


source: techcrunch ai: how ai guardrails are impeding the work of offensive cybersecurity researchers