model-releases
A 7% pass rate is not a step change. It is a benchmark with no floor yet.
GPT-6 Astra completed 7 of 100 dual-arm robotics tasks on a new benchmark, versus zero for a competing model. A researcher called it a step change. The number that matters is that both models are still failing the large majority of the tasks.
On a new robotics benchmark called StationeryBench, GPT-6 Astra completed 7 of 100 dual-arm manipulation tasks, according to The Decoder. A competing model, MolmoAct2, completed zero. A researcher quoted in the coverage called the result a "step change in spatial reasoning."
What the comparison actually shows
Seven successes against zero is a real gap, and worth reporting as one. But both numbers describe models that fail the task in the overwhelming majority of attempts. A benchmark where the best score is 7 percent has not yet established what a working system looks like; it has established that one model occasionally succeeds where another never does. "Step change" describes a benchmark that has moved from one working tier to a higher one. This benchmark has moved from zero working attempts to a small number of working attempts, on a scale that still rounds to failure.
Why the framing spreads anyway
A relative comparison, seven versus zero, is more dramatic to report than an absolute one, seven out of a hundred. Coverage built on the relative figure reads as validation of a capability leap; coverage built on the absolute figure reads as a benchmark still searching for its floor. Both are accurate. Only one is useful to a team deciding whether to build on the capability now.
A 7 percent pass rate is not evidence a capability has arrived. It is evidence the benchmark is hard enough to be worth watching.
What changes for anyone evaluating robotics claims
When a vendor or a benchmark report leads with a comparison to a competitor, ask for the absolute pass rate before treating the framing as a capability signal. A model that leads a field at 7 percent is not deployable for the task the benchmark measures, regardless of how it compares to the next-best model. Track the absolute number release over release; a genuine step change will show up as a jump from single digits into a range where a system succeeds more often than it fails, and StationeryBench is not there yet for anyone.
Sources
Written by the Devence Lab research team.