Insights

    model-releases

    A 7% pass rate is not a step change. It is a benchmark with no floor yet.

    Devence Lab

    · 2 min read

    Share
    Illustration · Devence Lab

    GPT-6 Astra completed 7 of 100 dual-arm robotics tasks on a new benchmark, versus zero for a competing model. A researcher called it a step change. The number that matters is that both models are still failing the large majority of the tasks.

    On a new robotics benchmark called StationeryBench, GPT-6 Astra completed 7 of 100 dual-arm manipulation tasks, according to The Decoder. A competing model, MolmoAct2, completed zero. A researcher quoted in the coverage called the result a "step change in spatial reasoning."

    What the comparison actually shows

    Seven successes against zero is a real gap, and worth reporting as one. But both numbers describe models that fail the task in the overwhelming majority of attempts. A benchmark where the best score is 7 percent has not yet established what a working system looks like; it has established that one model occasionally succeeds where another never does. "Step change" describes a benchmark that has moved from one working tier to a higher one. This benchmark has moved from zero working attempts to a small number of working attempts, on a scale that still rounds to failure.

    Why the framing spreads anyway

    A relative comparison, seven versus zero, is more dramatic to report than an absolute one, seven out of a hundred. Coverage built on the relative figure reads as validation of a capability leap; coverage built on the absolute figure reads as a benchmark still searching for its floor. Both are accurate. Only one is useful to a team deciding whether to build on the capability now.

    A 7 percent pass rate is not evidence a capability has arrived. It is evidence the benchmark is hard enough to be worth watching.

    What changes for anyone evaluating robotics claims

    When a vendor or a benchmark report leads with a comparison to a competitor, ask for the absolute pass rate before treating the framing as a capability signal. A model that leads a field at 7 percent is not deployable for the task the benchmark measures, regardless of how it compares to the next-best model. Track the absolute number release over release; a genuine step change will show up as a jump from single digits into a range where a system succeeds more often than it fails, and StationeryBench is not there yet for anyone.

    Sources

    1. GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarksThe Decoder

    Written by the Devence Lab research team.

    Share

    Collaborate

    We share findings with partners operating in the same constraint space.

    Get in touch