Wireva

OpenAI adjusts GPT-6 Astra benchmark figures after launch post goes live

OpenAI revised several evaluation metrics for its GPT-6 Astra model after publishing a delayed blog post, with some scores favoring Astra and others fluctuating for rival models.

This item was produced with AI assistance under the editorial responsibility of Haydamax OÜ.

OpenAI has revised several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement on Sept. 3, with some updated figures showing improved performance for Astra and weaker scores for models from rival Anthropic. The changes came amid an unusual rollout in which the post was briefly taken down and republished, and some numbers have continued to shift in the days since.

The company originally planned to publish the post at 2 p.m. ET, but it was not widely viewable until nearly two hours later. When OpenAI’s X account shared the link at 3:32 p.m., it returned an error message. At 3:50 p.m., CEO Sam Altman posted the link, writing that the company “hit a little snag getting the blog post deployed.” Multiple users, including reporters, were still unable to access it. OpenAI said the post was published shortly after 2 p.m. but retracted for reasons it would not disclose, first citing a content management system bug and then an internet outage.

Upon republication, the blog contained different evaluation metrics that appeared to favor Astra. One of the most notable changes involved Astra’s reported hallucination rate, which stood at 4.2% in early archived snapshots but was halved to 2% in a version captured at 5:20 p.m. The score for Astra’s predecessor, GPT-5.6 Sol, also dropped from 12.2% to 9.4%. As of this writing, both figures have reverted to their original levels.

OpenAI also boosted GPT-5.6 Sol’s score on its internal version of the ExploitBench cybersecurity evaluation from 5.5% to 11.5% in later versions. The company said it is investigating reverting that number because the higher result reflects a reasoning level not commercially available for Sol.

Mathematics scores also fluctuated. Astra’s FrontierMath Tier 4 score remained steady at 97.6%, but OpenAI briefly altered the scores for GPT-5.6 Sol and Anthropic’s Fable 5.1 model, making Astra appear significantly stronger in math. Anthropic’s score dropped from 87.8% to 78% before settling at 83%, while Sol’s went from 83% to 80.5% and back to 83%.

The metric adjustments began before the blog was first published. An embargoed draft provided to media outlets listed Astra’s score on the ARC-AGI-3 evaluation as 98.6%, but the live version now shows 99.99%. An OpenAI spokesperson said the company “always verify evals before publication” and that adjustments between draft and final versions are normal.

OpenAI noted that the Arc Prize Foundation, which created the ARC-AGI-3 benchmark, independently found Astra performed at 99.9% when given a particularly powerful harness, a set of tools the model can use to complete tasks. With the standard harness, Astra scored 63%, still better than any other publicly released AI model.

The episode highlights the intense competition in the AI industry and the challenges of measuring large language models with standardized benchmarks. Different research teams at OpenAI oversee various metrics and report them to a central team for publication. The company is open about the fact that scores are achieved under optimal conditions and may differ from the models available in the production ChatGPT product.

“Evaluation scores are the maximum at any effort,” reads a disclaimer on the blog. An OpenAI spokesperson said the company cares “deeply about getting evaluations right” and that most scores carry noise of a few percentage points depending on the checkpoint, scaffold, and evaluation run used. The company said it made fixes to ensure the numbers represent its best estimate of available model performance so users can make meaningful comparisons.

Same event, other desks

Story file →