Published on

GPT-6 Astra Rarely Refused Unsafe Robot Commands in New Test

Get your news fromHumanoids Daily

One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

P.A.
Written byP.A.

Supported by RoboStrategy

Sponsored

Investing involves risk, including possible loss of principal. Read the prospectus before investing.

  • RoboHarm reports that GPT-6 Astra attempted unsafe instructions in 97 of 100 robot-arm trials, completing 60.
  • Claude Fable 5.1 recorded 20 safety refusals, all in one scenario; Astra recorded two, plus one refusal unrelated to safety.
  • The findings concern five fixed laboratory scenes, not a general safety ranking or evidence of robots independently choosing to harm people.

GPT-6 Astra’s robotics results now come with an uncomfortable counterpart: a test of when it should stop.

In RoboHarm, Robocurve evaluated responses to unsafe instructions on robot arms. Astra attempted 97 of 100 trials and completed 60. The report was highlighted by co-author Jay Chooi on X.

The result adds a different dimension to our coverage of Astra outperforming specialist models on selected robotics tests. Better execution makes the question of which instructions a system will follow more consequential.

Refusal and completion measure different things

The report’s totals are:

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues
ModelSafety refusalsCompleted tasks
GPT-6 Astra2/10060/100
Claude Fable 5.120/10034/100
MolmoAct20/1006/100

Astra also refused once for a non-safety reason. Its 60 completions out of 97 attempts explain the approximately 62% figure in Chooi’s post; that is different from completion across all trials.

Fable’s safety refusals were concentrated entirely in the doll-and-knife scenario. Its aggregate refusal rate therefore masks a substantial difference between tasks.

What the experiment actually tested

The benchmark repository describes five fixed scenes involving a doll and knife, heating equipment, electrical objects, a battery and cleaning-product containers. Each model received one fixed instruction per scene across 20 independently reset trials. Reviewers assessed behavior using camera recordings and saved transcripts.

Some instructions referred indirectly to objects, requiring the system to interpret the scene. This was a test of instructed behavior, rather than robots developing harmful goals on their own. A doll stood in for a human-like target.

The repository also cautions that photographs do not establish every appliance’s power state or each container’s contents. Completing a scored movement should not automatically be described as causing an explosion, injury or toxic exposure.

A failed task is not a safety mechanism

The experiment’s labeling rubric separates refusal from an unsuccessful attempt. That distinction is essential: failing to manipulate an object does not demonstrate that a system recognized the danger and chose to avoid it.

It also changes how capability progress should be interpreted. If a system becomes better at carrying out instructions while its ability to reject unsafe ones stays unchanged, failures that previously prevented completion may disappear. That is an implication of the distinction, not a claim that this study establishes a universal relationship between intelligence and safety.

Our earlier reporting followed Astra’s block-placement results and a robot-arm painting demonstration. Those asked what a general model could accomplish through a robot. RoboHarm asks which requests it should decline.

The narrow setup leaves open how results would change with different wording, environments or additional safeguards. For deployment, the practical question is broader than whether a robot can finish a task: can the complete system reliably distinguish useful work from actions it should never execute?

Share this article

The weekly humanoid robotics briefing

One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.

Read recent issues