- Published on
OpenAI’s GPT-6 Astra Hits 95% on Physical Manipulation Benchmark, But Precision Tasks Expose Familiar Bottlenecks
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.


- In physical hardware trials conducted by RoboCurve, OpenAI’s GPT-6 Astra completed a block-placement task in 19 of 20 runs (95%), up from 40% for Claude Fable 5.1.
- Astra slashed token generation by 6.2x and cut run costs to $0.94, leveraging the architectural and spatial reasoning advances OpenAI unveiled in its flagship release.
- On a millimeter-tolerance puzzle insertion task, Astra hit the same wall as Claude Fable 5.1, stalling at a 10% completion rate (2/20).
- The tests were run using the open-source Inspect Robots framework on dual I2RT YAM arms, mapping multi-view camera inputs directly to 6-DoF end-effector targets.
- The results arrive as OpenAI prepares to build in-house humanoid hardware, underscoring both the rapid progress of LLM-driven kinematics and the enduring challenge of physical contact dynamics.
When OpenAI launched GPT-6 Astra, the lab highlighted breakthroughs across abstract reasoning, software engineering, and computer use. Among the benchmark records was a 95.9% geometric overlap score on BenchCAD—a test of generating 3D CAD structures directly from multi-view renders.
Now, independent hardware testing reveals that Astra’s spatial reasoning transfers directly to physical manipulation.
According to evaluation data released by public benefit corporation RoboCurve, GPT-6 Astra demonstrated dramatic leaps in spatial efficiency when tasked with controlling physical robot arms. Tested on dual I2RT YAM robotic arms via the open-source Inspect Robots framework, Astra completed a gross pick-and-place task with a 95% success rate, leaving Anthropic’s Claude Fable 5.1 (40%) and Claude Fable 5 (5%) far behind.
Yet the same trials also exposed the limits of zero-shot physical reasoning: as soon as a task demanded sub-millimeter insertion tolerances, Astra ran into the exact same wall that has confounded earlier foundation models.
Stay ahead in humanoid robotics
One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.
Faster Decisions, Fewer Tokens
RoboCurve tested the models across two physical setups: placing a red block into a bowl, and picking a circular wooden puzzle piece by its center knob to seat it into a matching groove.
In both benchmarks, the models acted as high-level agent policies under Inspect Robots. At each turn, the model ingested three camera feeds (an overhead perspective and two wrist-mounted views) alongside proprioceptive state data, emitting 6-DoF end-effector target poses that were passed to an inverse kinematics (IK) solver.
On the bowl task, Astra successfully placed the block in 19 out of 20 trials. Beyond the leap in reliability, Astra demonstrated a marked drop in reasoning overhead:
- Token Efficiency: Astra required an average of 2,100 output tokens per trial—a 6.2x reduction compared to Fable 5.1’s 12,900 tokens and an 89% reduction against Fable 5’s 19,200 tokens.
- Execution Latency: Runs averaged 2.5 minutes under Astra, down from 6.8 minutes for Fable 5.1 and 8.2 minutes for Fable 5.
- API Cost: At standard list pricing ($10 per million input tokens, $50 per million output tokens), Astra averaged $0.94 per completed run, compared to $2.12 for Fable 5.1.
The conciseness mirrors OpenAI's own findings in knowledge work, where Astra solved tasks with roughly 65% fewer output tokens than Claude Opus 5. Applied to physical kinematics, that efficiency translates to rapid decision loops. Jay Chooi of RoboCurve noted that if current performance curves hold, LLMs could be capable of controlling robotic arms in real time within two to three years.
The Millimeter Bottleneck
The dynamics changed completely during the precision puzzle task. Requiring the arms to grasp a slender knob and press a round piece into a tight circular recess, Astra completed just 2 of 20 trials (10%)—matching Claude Fable 5.1’s identical 10% mark, while Fable 5 failed entirely (0/20).
Under RoboCurve's human-scored five-stage rubric (from stage 0 for no approach to stage 4 for successful placement), Astra consistently reached stage 3, bringing the puzzle piece directly over the groove before hesitating or misaligning the final press.
While Astra still completed its attempts faster (3.4 minutes versus 5.9 minutes for Fable 5.1) and at a lower cost ($1.36 versus $2.18), the failure to finish the insertion reflects an industry-wide challenge.
As Unitree CEO Wang Xingxing recently argued when addressing why humanoids aren't yet working in factories, digital models operate in lossless vector spaces, whereas physical manipulation accumulates tiny physical errors with every movement. Without high-frequency tactile feedback and compliance, models struggle to correct for micro-deviations in the final millimeters of contact.
Standardized Auditing with Inspect Robots
The benchmark also marks an important milestone for Inspect Robots, RoboCurve’s open-source framework for physical AI evaluation.
Modeled on the UK AI Safety Institute’s Inspect AI platform, the library provides a unified interface for testing foundation policies—including frontier LLMs, code-as-policy agents, and served vision-language-action (VLA) models—against simulators like Isaac Lab or physical rigs from Franka, Unitree, and AgiBot. The system logs synchronized multi-camera streams, joint trajectories, and tool-call transcripts, generating replayable visualizations via Rerun.
RoboCurve noted several methodological constraints in its initial trial release:
- Hardware Allocation: While all puzzle trials ran on the same physical setup (rig-4), the bowl evaluations placed Astra on rig-1 due to maintenance on rig-3, where the Fable runs took place.
- Non-Interleaved Runs: Astra's trials were executed two days after the Anthropic runs rather than interleaved sequentially.
- Operator Scoring: Human observers graded stage progression with knowledge of which model was active, leaving room for scoring bias.
- Prompt Caching: Calculations used standard list prices; OpenAI automatically cached roughly 20% of Astra’s input prompts, meaning actual operational expenses for Astra were likely lower than reported.
Closing the Physical Loop
The emergence of Astra as a capable kinematic planner arrives at a pivotal moment for OpenAI. CEO Sam Altman recently confirmed that the company will "definitely" build its own humanoid robot, settling months of speculation following the launch of its dedicated internal robotics division.
That hardware ambition has already disrupted OpenAI's previous alliances. Startups like Figure AI severed joint development efforts to pursue in-house foundation architectures like Helix, backing up their efforts with crowdsourced real-world data platforms like Index. Meanwhile, hardware giants like NVIDIA continue to anchor the open-source pipeline, recently acquiring Hugging Face for $12.9 billion to bridge physical simulation and open model hubs.
RoboCurve’s evaluation demonstrates why frontier AI developers are eager to step into physical robotics. When standard reasoning models can handle 95% of gross manipulation out of the box, building physical embodiments becomes the logical next frontier.
Yet until foundation models can perceive tactile resistance and adapt to dynamic contact, that last millimeter will remain the hardest problem on the bench.
Share this article
Read next
- Published on
- Reading time
- 5 min read
Sam Altman Confirms OpenAI "Will Definitely Do a Humanoid"
- Published on
- Reading time
- 4 min read
OpenAI Pivots Directly Into Hardware, Launching Internal Robotics Division
- Published on
- Reading time
- 3 min read
OpenAI Hardware Leader Caitlin Kalinowski Resigns Over Pentagon Deal as Benjamin Bolte Joins
- Published on
- Reading time
- 7 min read
NVIDIA Buys Hugging Face for $12.9B as Open-Source Robotics Scales Up
- Published on
- Reading time
- 5 min read
Figure Inks Multi-Billion-Dollar Compute Deal With Nscale to Deploy 100,000 Next-Gen NVIDIA GPUs
- Published on
- Reading time
- 7 min read
The $1,688 Appliance Bet: Nori Robotics Launches Wheeled Bimanual Manipulator
Stay ahead in humanoid robotics
One email a week: the launches, funding and research that mattered. Join the researchers, founders and technologists following the field.












