GPT-6 Astra benchmark changes made by OpenAI after the model’s launch on 3 September 2026 included shifting hallucination rates, adjusted rival scores and a cybersecurity figure the company says it is now considering reverting.
The figures on OpenAI’s announcement blog shifted multiple times in the hours after publication, with some metrics moving in Astra’s favour and others changing back. ExplainX confirmed the launch date as 3 September 2026.
Hallucination rate swung from 4.2% to 2% and back
The first internet archive snapshot of the blog, taken at 2:23pm ET, listed Astra’s hallucination rate at 4.2%. By a later snapshot at 5:20pm, that figure had been halved to 2%. As of the time of writing, it has returned to 4.2%.
The hallucination rate for Astra’s predecessor, GPT-5.6 Sol, also shifted, dropping from 12.2% to 9.4% before returning to 12.2%.
OpenAI’s spokesperson told Fortune: ‘We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting.’
The company said it made ‘fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.’
GPT-6 Astra benchmark changes extended to rival model scores
Scores for Anthropic‘s Fable 5.1 model on the maths evaluation FrontierMath Tier 4 (v2) dropped from 87.8% in the first snapshot to 78% by 5:17pm, before recovering to 83%. GPT-5.6 Sol’s score on the same test went from 83% to 80.5% and back to 83%. Astra’s own score on that evaluation stayed at 97.6% throughout.
On OpenAI’s internal ExploitBench cybersecurity evaluation, Sol’s score jumped from 5.5% in the first version to 11.5% in later versions. OpenAI said it is currently investigating reverting that number to 5.5%, saying the 11.5% result reflects a reasoning level that is not commercially available for Sol.
Not every change favoured OpenAI’s models. Anthropic’s Claude Fable 5.1 score on the healthcare evaluation HealthBench Professional improved from 56.6% to 58.1% in revised versions, and Opus 5’s score rose from 54.5% to 56.4%.
Chaotic rollout preceded the revisions
OpenAI originally planned the blog post to go live at 2pm ET. It was briefly published then retracted; the company gave Fortune two conflicting explanations, first saying a bug in the content management system was to blame, then citing an internet outage. Sam Altman posted the link at 3:50pm, writing: ‘We hit a little snag getting the blog post deployed, but it is really great.’
When OpenAI republished the blog, the evaluation metrics had changed. Some figures continued to change after that.
An embargoed pre-publication draft sent to media had listed Astra’s ARC-AGI-3 score at 98.6%. The live blog now shows 99.99%. OpenAI said ‘adjustments between draft and final version are normal’ as evaluations are verified before publication.
Industry experts question ‘benchmaxxing’
Researchers Anka Reuel and Mike Hardy of the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab said the practice of re-running evaluations under different conditions, known as ‘benchmaxxing’, ‘can be done in a very tight timeframe, and it’s better for their marketing.’
They also said Astra’s system card ‘barely any details about the evaluation’ for the internal hallucination benchmark, including no count of test items.
Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, said shifting scores before a launch are ‘usually a function of final launch logistics’, as checkpoints, configurations and harnesses are ‘typically still shifting in the final days before a launch.’ He called for industry norms requiring companies to disclose what changed when benchmark numbers are revised.
OpenAI’s blog carries a disclaimer that ‘evaluation scores are the maximum at any effort’, with further caveats in footnotes on each metric. The company said ‘things like harness, reasoning level and other factors inform evals.’
OpenAI said it is investigating reverting Sol’s ExploitBench figure to 5.5%.

