Insights

    Model Releases

    Capability thresholds are becoming a disclosure norm. Deployers should read them as a handoff.

    Devence Lab

    · 2 min read

    Share
    Capability thresholds are becoming a disclosure norm. Deployers should read them as a handoff.
    Photograph · Unsplash

    Labs now publish where they think a model crosses into dangerous capability. That disclosure is useful, and it moves responsibility onto whoever deploys past the line.

    Across recent frontier releases, a pattern has settled: labs publish an assessment of whether a model crosses defined capability thresholds, particularly in cyber. One model is described as meeting a critical cybersecurity threshold; another is characterised through refusal rates against malicious agentic requests.

    This is a genuine improvement in transparency. It is also, read carefully, a transfer.

    What a threshold disclosure does

    It establishes on the record, with a date, that the vendor believed this system could materially assist a capable attacker. Whatever else follows, nobody can later claim the risk was unknown or unstated.

    For the lab, that is the point — it discharges a duty to warn and defines the boundary of their responsibility. For the deployer, it means proceeding is now a documented decision to operate a system publicly characterised as dangerous. The disclosure did not make the model riskier; it made your decision legible.

    A published threshold converts an unknown risk into an accepted one, and acceptance has an owner.

    Using it properly

    The useful response is not to avoid threshold-crossing models. It is to let the disclosure drive the control set. If a vendor says a model can autonomously discover exploitable vulnerabilities, that is a specification for what your monitoring needs to catch and what your authorisation layer needs to prevent.

    Concretely: the model's own capability profile tells you what a compromised or misdirected instance would be able to do in your environment. That is a threat model handed to you for free, and most organisations file it rather than use it.

    Where the disclosures fall short

    Thresholds are assessed against the lab's evaluation set, in the lab's deployment context, with the lab's safeguards active. Your context has different tools, different data and different adversaries.

    A model below a threshold in isolation can be well past it once wired to your systems, because capability is a property of the assembled system rather than the weights. No vendor disclosure covers that, and it is precisely the gap your own assurance work has to fill.

    Sources

    1. Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access ProgramsThe Hacker News
    2. EU Artificial Intelligence Act — developments and analysesartificialintelligenceact.eu
    3. AI Updates Today (September 2026) — Latest AI Model Releasesllm-stats.com

    Written by the Devence Lab research team.

    Share

    Collaborate

    We share findings with partners operating in the same constraint space.

    Get in touch