AI Security
Three labs shipped cyber models in one week. The capability is not the story.
Google, Anthropic and OpenAI all put offensive-capable security models behind access programmes in early September. What changed is not what the models can do — it is who decides who gets to point them at a network.
In the first week of September, three frontier labs shipped cyber-capable models within days of each other. Google released Gemini 3.8 Flash Cyber alongside a Fairwind access programme for what it calls high-priority defenders. Anthropic launched Claude Mythos 5.1 with an Enterprise Frontier Safeguards tier pairing zero data retention with its own guardrails. OpenAI's Astra was published against a Critical cybersecurity capability threshold, reported at 100% on ExploitBench and declining 91.5% of jailbreaking attempts, distributed through a programme called Daybreak Blue.
The coverage has focused on the benchmarks. That is the least interesting part of the week.
Gated distribution is the actual product decision
Every one of these releases arrived wrapped in an access programme. Not an API key and a credit card — an application, a review, a partner list. Google's programme names 650-plus partners including CrowdStrike, Datadog, Palo Alto Networks and Snowflake. That is not a safety afterthought bolted onto a launch. It is the launch.
For anyone running a security programme, this reframes a procurement question that used to be simple. You are no longer buying a capability that you then govern. You are entering a relationship in which the vendor governs your access to it, continuously, and can revoke it. The model is the easy part. The eligibility criteria, the revocation terms and the audit obligations attached to them are what your legal team should be reading.
A capability you can lose on someone else's judgement is not a capability you can build a control around.
The threshold language is a liability boundary
OpenAI describing Astra as meeting a Critical cybersecurity capability threshold is not marketing. It is a declaration, made in public, that the vendor believes this system can materially assist an attacker. Anthropic's framing around refusal rates does similar work from the other direction.
Read those statements as what they are: an attempt to establish where the vendor's responsibility ends and the operator's begins. If a model is publicly documented as critical-capability and you deploy it without commensurate controls, the disclosure is already on the record. That matters less for the labs than for the enterprises. The threshold language will show up in incident post-mortems long before it shows up in regulation.
A defender-only model is a category error
The framing across all three launches is defensive: find your own vulnerabilities before someone else does. The capability is symmetric. A system that discovers exploitable flaws autonomously does so regardless of the intent of the operator holding the key, and the access programme is the only thing standing between those two populations.
That is a reasonable control. It is also a single point of failure, and it is administrative rather than technical. The interesting question for the next twelve months is not whether these models are good at finding bugs. It is what happens the first time a legitimately enrolled partner is itself compromised, and the attacker inherits the enrolment.
What to do about it this quarter
Three concrete things. First, inventory which of your vendors have enrolled in these programmes — your attack surface now includes their eligibility. Second, treat model access credentials as privileged infrastructure credentials, not as API keys; they belong in the same rotation and monitoring regime as your cloud root accounts. Third, write down what you would do if access were revoked mid-engagement, because the terms permit it and nobody has rehearsed it.
The capability arrived this month. The governance for it has not, and the access programmes are a placeholder standing in for controls that do not exist yet.
Sources
Written by the Devence Lab research team.