Published on

GPT-6 Astra Leads SuperCLUE’s Robot-Brain Benchmark, With a Wide Planning Advantage

Get your news fromHumanoids Daily

One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Humanoids Daily
Written byHumanoids Daily

OpenAI’s GPT-6 Astra posts the highest overall score among the ten models in SuperCLUE’s September 2026 EmbodiedCLUE-VLA “Embodied Brain” evaluation. Its clearest advantage is in interaction and planning: 85.71, compared with runner-up Gemini-3.8-Flash’s 64.29.

  • Astra leads the displayed models overall, with a 21.42-point advantage in interaction and planning over Gemini-3.8-Flash.
  • International models appear as references, outside SuperCLUE’s numbered rankings.
  • These scores assess cognitive capabilities, not demonstrated robot task success.
SuperCLUE’s September 2026 embodied-brain overall-score chart for ten models, led by GPT-6 Astra at 86.8 and Gemini-3.8-Flash at 82.1.
September 2026 overall scores, rounded to one decimal place. International models are displayed for reference outside the numbered rankings. Chart: SuperCLUE.

The ten-model evaluation covers basic perception, visual reasoning, interaction and planning, and embodied safety. Selected results show how much the overall score can conceal:

ModelOverallInteraction & planning
GPT-6 Astra86.7585.71
Gemini-3.8-Flash82.1264.29
Qwen3.8-Max-090277.4839.29
Nebula-EmbodiedBrain77.4839.29
Gemini-Robotics-ER-2-Preview70.8646.43
MiMo-Embodied-7B56.9514.29

Source: SuperCLUE, September 2026. International models are included for reference without numbered ranks.

Astra’s overall lead over Gemini-3.8-Flash is 4.63 points. The planning gap is considerably wider. Qwen and Nebula share the top Chinese rank, while the robotics-focused Gemini-Robotics-ER-2-Preview also trails Astra on both measures.

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues

The tested field also needs context. In its September report, SuperCLUE says it selected ten representative domestic and international models, including five with no more than 10 billion parameters. It does not explain the inclusion criteria further. Astra’s lead therefore applies to this tested group.

The evaluation uses images and Chinese prompts, with answers scored against references through judge models or rule-based scripts. Crucially, every model receives a fixed input/output protocol without model-specific adaptation. SuperCLUE acknowledges that specialist models may depend on particular prompts, reasoning modes or action-output formats. It says their low planning scores partly reflect this protocol choice and interprets the results as exposing limitations in generalization across tasks and protocols.

That makes the comparison useful, but also narrower than a verdict on each model’s best possible robotics performance. A standardized protocol can test how readily models handle a common interface while missing capabilities unlocked by their intended configuration.

For robotics developers, that breakdown is more useful than a single league-table position. Recognizing objects, understanding a scene and choosing a workable sequence of actions are different demands. A strong aggregate score can obscure a weakness in the capability a particular application needs most.

SuperCLUE describes this benchmark as assessing embodied agents’ cognitive core. The results should therefore inform questions about robot intelligence, rather than settle purchasing or deployment decisions. A safety-category score is not a safety certification, and a planning score does not establish that a robot can reliably execute a plan with its own hardware.

The next question is how these differences hold up when models must act, observe the consequences and recover from mistakes. That would require evidence from physical trials beyond this leaderboard.

Share this article

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues