---
title: "A 7% pass rate is not a step change. It is a benchmark with no floor yet."
description: "GPT-6 Astra completed 7 of 100 dual-arm robotics tasks on a new benchmark, versus zero for a competing model. A researcher called it a step change. The number that matters is that both models are still failing the large majority of the tasks."
url: "https://devencelab.com/insights/2026/09/13/a-7-pass-rate-is-not-a-step-change"
date: "2026-09-13"
section: "Insights"
tag: "Model Releases"
author: "Devence Lab"
reading_time: "2 min read"
site: "Devence Lab"
license: "Readable and quotable with attribution to the canonical URL."
---

# A 7% pass rate is not a step change. It is a benchmark with no floor yet.

GPT-6 Astra completed 7 of 100 dual-arm robotics tasks on a new benchmark, versus zero for a competing model. A researcher called it a step change. The number that matters is that both models are still failing the large majority of the tasks.

On a new robotics benchmark called StationeryBench, GPT-6 Astra completed 7 of 100 dual-arm manipulation tasks, according to The Decoder. A competing model, MolmoAct2, completed zero. A researcher quoted in the coverage called the result a "step change in spatial reasoning."

## What the comparison actually shows

Seven successes against zero is a real gap, and worth reporting as one. But both numbers describe models that fail the task in the overwhelming majority of attempts. A benchmark where the best score is 7 percent has not yet established what a working system looks like; it has established that one model occasionally succeeds where another never does. "Step change" describes a benchmark that has moved from one working tier to a higher one. This benchmark has moved from zero working attempts to a small number of working attempts, on a scale that still rounds to failure.

## Why the framing spreads anyway

A relative comparison, seven versus zero, is more dramatic to report than an absolute one, seven out of a hundred. Coverage built on the relative figure reads as validation of a capability leap; coverage built on the absolute figure reads as a benchmark still searching for its floor. Both are accurate. Only one is useful to a team deciding whether to build on the capability now.

> A 7 percent pass rate is not evidence a capability has arrived. It is evidence the benchmark is hard enough to be worth watching.

## What changes for anyone evaluating robotics claims

When a vendor or a benchmark report leads with a comparison to a competitor, ask for the absolute pass rate before treating the framing as a capability signal. A model that leads a field at 7 percent is not deployable for the task the benchmark measures, regardless of how it compares to the next-best model. Track the absolute number release over release; a genuine step change will show up as a jump from single digits into a range where a system succeeds more often than it fails, and StationeryBench is not there yet for anyone.

## Sources

- [GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks](https://the-decoder.com/gpt-6-astra-appears-to-show-a-step-change-in-spatial-reasoning-based-on-early-benchmarks/) - The Decoder
