Open-Weight AI Models Close the Cyber Capability Gap to Four Months, UK Institute Finds
The UK's AI Security Institute says open-weight models now trail the best closed systems by just four to seven months on cyber tasks — and match them at a fraction of the cost, a 'persistent and irreversible risk of misuse.'
In Brief
- The UK AISI finds open-weight models now trail closed frontier models by four to seven months, down from six to ten.
- GLM-5.2 matched Anthropic’s Opus 4.6 and DeepSeek V4-Pro matched Opus 4.5 on cyber tasks.
- A full 100M-token cyber test cost $85 on Opus versus just $1.19 on DeepSeek V4-Pro.
Open-weight AI models have closed the gap with the most capable closed systems on cyber operations to roughly four to seven months, down from six to ten months, according to a new assessment from the UK’s AI Security Institute (AISI). The finding arrives as open models also match frontier performance at a fraction of the cost.
In the institute’s testing, Zhipu’s GLM-5.2, released in June 2026, matched the cyber performance of Anthropic’s Opus 4.6 from February 2026 — about four months newer — while DeepSeek V4-Pro matched Opus 4.5, AISI said.
The capability is no longer just the domain of well-funded labs: a full 100-million-token cyber test cost about $85 on Opus 4.6 but only $1.19 on DeepSeek V4-Pro, and about $46 on GLM-5.2 — a cost collapse the institute called a “persistent and irreversible risk of misuse.”
How open-weight models now perform on cyber tasks
AISI evaluated models on “Narrow Cyber Tasks,” a set of roughly 70 tasks graded at four difficulty levels, and on “Cyber Ranges” — complex, multi-step scenarios. The Last Ones range spanned 32 steps across four subnets and about 20 hosts, work that took a human about 20 hours.
The headline result is the compression of the lag: open models that once trailed by half a year or more now sit within months of the frontier, and on specific tasks they are at parity, eroding the safety buffer that closed development once provided.
Cost is the force multiplier. At under two dollars for a full benchmark run, open-weight models put frontier-grade cyber assistance within reach of almost any actor, a shift the UK’s National Cyber Security Centre had already flagged in warnings issued in April 2026.
What the cost collapse means for AI safety
AISI’s verdict is blunt: the combination of near-frontier capability and near-free cost creates a “persistent and irreversible risk of misuse” that cannot be undone by further restricting the strongest closed models.
The institute’s data shows the per-task economics: Opus 4.6 cost about $15 per task, GLM-5.2 about $6, and DeepSeek V4-Pro about $0.28 — a more than 50-fold difference that reshapes who can run sustained offensive operations.
The report lands as policymakers weigh how to govern open releasing of weights, with the UK framing the finding as evidence that the debate can no longer assume capability stays concentrated behind paywalls and enterprise contracts.
The UK AI Security Institute found open models now trail closed systems by four to seven months on cyber tasks, down from six to ten, in an assessment reported by THE DECODER.
FAQ
How far behind are open-weight models now?
About four to seven months behind closed frontier models, down from six to ten, per the UK AI Security Institute.
Which open models matched frontier cyber performance?
Zhipu’s GLM-5.2 matched Opus 4.6 and DeepSeek V4-Pro matched Opus 4.5 on cyber tasks.
Why is the cost drop a problem?
A full cyber benchmark run cost about $1.19 on DeepSeek V4-Pro versus $85 on Opus, a gap AISI called a ‘persistent and irreversible risk of misuse.’