Model Releases
7 analyses from Devence Lab
· 2 min read
A 7% pass rate is not a step change. It is a benchmark with no floor yet.
GPT-6 Astra completed 7 of 100 dual-arm robotics tasks on a new benchmark, versus zero for a competing model. A researcher called it a step change. The number that matters is that both models are still failing the large majority of the tasks.
· 2 min read
Release notes just became compliance artifacts
· 2 min read
Capability thresholds are becoming a disclosure norm. Deployers should read them as a handoff.
· 2 min read
Your model's deprecation date is a risk you do not control
· 2 min read
A model scored 100% on ExploitBench. That tells you about the benchmark.
· 2 min read
GPT-6, Grok 4.7 and Gemini 3.8 shipped inside ten days. Your qualification cycle did not.
· 2 min read
The first cyber-defence model shipped as a product, not a research artefact