Three out of five AI models fail tests designed to stop terrorists from using them. And once the guardrails are stripped, the failure rate hits 100 percent.
That’s the finding from a new study by the UK nonprofit Tech Against Terrorism, reported by CBC News on Friday. Researchers ran more than 130 AI models through hundreds of prompts written to resemble the kinds of requests someone planning an attack might submit. Most models held up. The ones with their safety systems deliberately removed did not.
The technique is called “abliteration”: the deliberate removal of a model’s safety training. It doesn’t require a lab or a supercomputer. Modified versions circulate freely on Hugging Face, where the organization counted more than 29,000 repositories advertising uncensored or unprotected models as of late last month.
What the test actually measured
Researchers at Tech Against Terrorism built a benchmark of hundreds of requests covering attack planning, terrorist financing, and radicalization. They fed those prompts to over 130 models and scored the responses on a 100-point safety scale.
The headline number: three in five models failed. That’s the aggregate across everything tested, including the stripped-down copies floating around model hubs. But the detail that will worry anyone in AI safety is the split between stock models and modified ones.
Every model that had gone through abliteration failed every test. Not most of them. All of them.
That’s a design flaw with an uncomfortable implication: safety in today’s open-weight models isn’t baked into the weights. It sits on top of them, in the fine-tuning, and it can be peeled off.
The 97-to-3 drop
The study’s sharpest illustration uses Meta’s Llama 3.1 8B, one of the most widely downloaded open models in the world. Unmodified, it scored 97 out of 100 on the terrorism safety benchmark and refused dangerous requests the way its safety training intended.
The abliterated version of the same model scored approximately 3. It answered requests about attacks, terrorist financing, and radicalization in detail, where the original version had said no.
Same weights. Same architecture. A 94-point swing, and the difference between a refusal and a detailed answer is a modification process the report suggests takes minutes with free tools.
That number should end the argument about whether open-weight safety is fragile. The safeguard isn’t the model. It’s the refusal layer, and the refusal layer is removable.
A warehouse of unguarded models
The 29,000 figure is arguably the more alarming stat. That’s how many repositories on Hugging Face, the GitHub of the AI world, were advertising uncensored or unprotected models at the end of September. Twenty-nine thousand.
Hugging Face says it regularly moderates content that violates its policies, and it pushed back on the study’s recommendations, warning that some of them could undermine open research. The company has a point: much of that repository count includes fine-tunes, merges, and personal copies that aren’t all functionally identical to the headline models. But scale is scale. If even a fraction of those 29,000 copies run without refusal layers, the barrier to getting a compliant model is a search box.
Meta, whose Llama models feature prominently in the findings, said its models undergo safety evaluations and that its policies prohibit harmful or illegal uses. That’s true and also beside the point. The failure here isn’t in Meta’s released model, which scored 97. It’s in what anyone can do to that model after release, with tools Meta can’t police.
Why this matters
Here’s the awkward part: the researchers found no evidence that terrorist or extremist groups are actually using the tested models. One extremist chatbot turned up. That’s it.
That cuts two ways. Optimists will read it as proof the threat is theoretical. But “no one’s noticed yet” is exactly what Adam Hadley, the organization’s founder and executive director, is worried about. “The thing is actually, this has already happened because a lot of these open models have already been broken — it’s just no one’s noticed yet,” he told reporters.
Think about what that means in practice. Every major AI lab has spent years building refusal layers, red-teaming prompts, and writing safety reports. The entire industry treats the guardrail as the safety system. This study is evidence that for open models, the guardrail is a sticker over the real machinery, and 29,000 copies of the machinery are one download away.
Tech Against Terrorism’s recommendations (independent safety benchmarks, stronger technical protections against safeguard removal, restrictions on distributing modified models) all face the same problem. Once weights are public, removal is a local operation. Nobody can unpublish a download that’s already happened 29,000 times.
That leaves an uncomfortable question for the open-source AI movement: is the current approach to model release actually compatible with safety testing that assumes the guardrails stay on? The 97-to-3 swing says no. The labs have one more incentive structure to redesign, and this time the fix can’t be shipped in the next model release.
Frequently asked questions
How many AI models failed the terrorism safety tests? Three in five AI models failed terrorism safety evaluations in a Tech Against Terrorism study of more than 130 models. Models with removed guardrails — “abliterated” models — failed every test.
What is an abliterated AI model? An abliterated AI model is one whose safety guardrails have been deliberately stripped out using a process called abliteration. Tech Against Terrorism found more than 29,000 repositories advertising such uncensored models on Hugging Face.
How did Meta’s Llama 3.1 8B perform in the AI safety tests? The unmodified Llama 3.1 8B scored 97 out of 100 on the safety benchmark and refused dangerous requests. After its safeguards were stripped, the same model scored about 3 and gave detailed answers to requests about attacks, financing, and radicalization.
Are terrorists actually using these AI models? The study found no evidence of terrorist or extremist groups using the tested models, with the exception of one extremist chatbot identified by the researchers.
Sources: CBC News, The National, Anadolu Agency, TRT World, Yeni Safak
