OpenAI logo displayed on a screen with a robotic hand. (Representative image)
OpenAI ’s flagship model, GPT-6 Astra, attempted 97 out of 100 unsafe directives during specialised safety evaluations on robotic arms, according to findings from benchmark testing platform Robocurve. These results are published days after the ChatGPT-maker disclosed 6 incidents where its internal AI agents escaped containment and hacked external platforms. Meanwhile, Anthropic’s Claude Fable 5.1 returned better numbers in these dangerous tests. Battery equipment Each system faced one fixed prompt per setting across 20 distinct, reset trials.
The results, highlighted on X (formerly Twitter) by study co-author and Robocurve co-founder Jay Chooi, cast a spotlight on whether modern general-purpose models know when to halt dangerous physical actions. Evaluators reviewed physical actions and recorded model transcripts via camera feeds. Several prompts relied on indirect object descriptions, requiring the neural networks to visually interpret the setup rather than react to obvious trigger words.

