All news

AISI: open-weight models trail frontier on cyber by 4–7 months

AI Security Institute (UK)Summarized by the ORA·tech AI assistant
AISI: open-weight models trail frontier on cyber by 4–7 months

The UK AI Security Institute published its first public analysis of the cyber capability gap between open-weight and closed frontier models: GLM-5.2 and DeepSeek V4-Pro trail by just 4–7 months, down from 6–10 months in 2025.

The UK AI Security Institute (AISI), part of the Department for Science, Innovation and Technology, published on July 17, 2026 its first public analysis of how far open-weight models trail the closed-weight frontier on cyber capability. The finding: the gap is narrowing.

Key points

  • GLM-5.2 (June 2026) was the most cyber-capable open-weight model at time of testing, comparable to Opus 4.6 (Feb 2026) on AISI's narrow cyber tasks and to Opus 4.5 (Nov 2025) on its cyber ranges — a 4 to 7 month lag.
  • DeepSeek V4-Pro is comparable to Opus 4.5, released 5 months earlier. Both gaps are narrower than the 6–10 months AISI measured internally through most of 2025.
  • Testing spanned 70 narrow cyber tasks across four difficulty tiers, plus the "The Last Ones" cyber range — a 32-step corporate network attack across 4 subnets and roughly 20 hosts, estimated at ~20 hours for a human expert.
  • Cost gap is large: a 100M-token cyber range run cost roughly $85 for Opus 4.5 and 4.6, versus an estimated $46 for GLM-5.2 and $1.19 for DeepSeek V4-Pro.
  • Safeguards barely slowed testing: DeepSeek V4-Pro occasionally refused reverse-engineering tasks, which AISI circumvented with a small number of repeat attempts.
  • AISI says it intends to evaluate Kimi K3 on the same basis once its weights are released, announced for end of July.
Source
AI Security Institute (UK)
Read the original
#AI security#open weight#LLM#cyber#AISI
This summary was written by the ORA·tech AI assistant. Read the original for full context.

Related