Close Menu
Daily NewsDaily News
    Pages
    • Home
    • About
    • Meet the Daily News Team
    • Contact
    • Terms and Conditions
    • Privacy Policy
    Facebook X (Twitter) Instagram
    Facebook X (Twitter)
    Daily NewsDaily News
    Subscribe
    • News
    • Entertainment
    • Finance
    • Health
    • Lifestyle
    • UK Politics
    • Property
    • Technology
    • Travel
    • World
    Daily NewsDaily News
    Home » Latest » OpenAI altered GPT-6 Astra benchmark scores after chaotic launch
    News

    OpenAI altered GPT-6 Astra benchmark scores after chaotic launch

    Philip MarchettiBy Philip Marchetti19/09/20264 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email Reddit WhatsApp Copy Link
    GPT-6 Astra benchmark changes
    Share
    Facebook Twitter LinkedIn Pinterest Email Reddit WhatsApp Copy Link

    GPT-6 Astra benchmark changes made by OpenAI after the model’s launch on 3 September 2026 included shifting hallucination rates, adjusted rival scores and a cybersecurity figure the company says it is now considering reverting.

    The figures on OpenAI’s announcement blog shifted multiple times in the hours after publication, with some metrics moving in Astra’s favour and others changing back. ExplainX confirmed the launch date as 3 September 2026.

    Hallucination rate swung from 4.2% to 2% and back

    The first internet archive snapshot of the blog, taken at 2:23pm ET, listed Astra’s hallucination rate at 4.2%. By a later snapshot at 5:20pm, that figure had been halved to 2%. As of the time of writing, it has returned to 4.2%.

    The hallucination rate for Astra’s predecessor, GPT-5.6 Sol, also shifted, dropping from 12.2% to 9.4% before returning to 12.2%.

    OpenAI’s spokesperson told Fortune: ‘We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting.’

    The company said it made ‘fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.’

    GPT-6 Astra benchmark changes extended to rival model scores

    Scores for Anthropic‘s Fable 5.1 model on the maths evaluation FrontierMath Tier 4 (v2) dropped from 87.8% in the first snapshot to 78% by 5:17pm, before recovering to 83%. GPT-5.6 Sol’s score on the same test went from 83% to 80.5% and back to 83%. Astra’s own score on that evaluation stayed at 97.6% throughout.

    On OpenAI’s internal ExploitBench cybersecurity evaluation, Sol’s score jumped from 5.5% in the first version to 11.5% in later versions. OpenAI said it is currently investigating reverting that number to 5.5%, saying the 11.5% result reflects a reasoning level that is not commercially available for Sol.

    Not every change favoured OpenAI’s models. Anthropic’s Claude Fable 5.1 score on the healthcare evaluation HealthBench Professional improved from 56.6% to 58.1% in revised versions, and Opus 5’s score rose from 54.5% to 56.4%.

    Chaotic rollout preceded the revisions

    OpenAI originally planned the blog post to go live at 2pm ET. It was briefly published then retracted; the company gave Fortune two conflicting explanations, first saying a bug in the content management system was to blame, then citing an internet outage. Sam Altman posted the link at 3:50pm, writing: ‘We hit a little snag getting the blog post deployed, but it is really great.’

    When OpenAI republished the blog, the evaluation metrics had changed. Some figures continued to change after that.

    An embargoed pre-publication draft sent to media had listed Astra’s ARC-AGI-3 score at 98.6%. The live blog now shows 99.99%. OpenAI said ‘adjustments between draft and final version are normal’ as evaluations are verified before publication.

    Industry experts question ‘benchmaxxing’

    Researchers Anka Reuel and Mike Hardy of the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab said the practice of re-running evaluations under different conditions, known as ‘benchmaxxing’, ‘can be done in a very tight timeframe, and it’s better for their marketing.’

    They also said Astra’s system card ‘barely any details about the evaluation’ for the internal hallucination benchmark, including no count of test items.

    Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, said shifting scores before a launch are ‘usually a function of final launch logistics’, as checkpoints, configurations and harnesses are ‘typically still shifting in the final days before a launch.’ He called for industry norms requiring companies to disclose what changed when benchmark numbers are revised.

    OpenAI’s blog carries a disclaimer that ‘evaluation scores are the maximum at any effort’, with further caveats in footnotes on each metric. The company said ‘things like harness, reasoning level and other factors inform evals.’

    OpenAI said it is investigating reverting Sol’s ExploitBench figure to 5.5%.

    Post Views: 5
    Follow on Google News Follow on Facebook Follow on X (Twitter)
    Share. Facebook Twitter LinkedIn Tumblr Email Reddit WhatsApp Copy Link
    Previous ArticleThrive Capital FIFA regrets deepen as UEFA files in three US courts
    Philip Marchetti

    Philip Marchetti spent a decade in broadcast journalism before moving to print and digital. He started as a researcher at a regional TV newsroom, worked his way onto the news desk, and spent five years producing packages on everything from council corruption to factory closures across the Midlands. He went freelance in 2019 and started writing because he missed the reporting and did not miss the rota. He covers UK politics, public services, and the slow-moving institutional stories that only make the front page when something breaks. Philip lives in Nottingham. He reads select committee transcripts the way other people read thrillers, and finds them roughly as plausible.

    Related Posts

    By Philip Marchetti18/09/2026

    Thrive Capital FIFA regrets deepen as UEFA files in three US courts

    By Philip Marchetti18/09/2026

    Women dominate US job gains, taking 98% of August’s 162,000 new posts

    By Philip Marchetti18/09/2026

    Nevada Developer in Talks Over Yosemite Land Swap Deal for Private Road

    Top Stories

    OpenAI altered GPT-6 Astra benchmark scores after chaotic launch

    19/09/2026

    Thrive Capital FIFA regrets deepen as UEFA files in three US courts

    18/09/2026

    Women dominate US job gains, taking 98% of August’s 162,000 new posts

    18/09/2026

    Nevada Developer in Talks Over Yosemite Land Swap Deal for Private Road

    18/09/2026
    Topics
    • Accessories
    • Adventure
    • Aerospace & Defence
    • Animal
    • Animals & Pets
    • Art & Culture
    • Automotive
    • Awards
    • Banking
    • Books & Publishing
    • Business
    • Business & Retail
    • Career
    • Charity
    • Community
    • Culture & Art
    • Cybersecurity
    • Defence
    • Design & Innovation.
    • Economics
    • Economy
    • Education
    • Electronics
    • Employment
    • Energy
    • Entertainment
    • Environment
    • Event
    • Events & Festivals
    • Fashion
    • Fashion & Beauty
    • Featured
    • Festivals
    • Finance
    • food
    • Food & Beverage
    • Gaming
    • Health
    • Homes & Interiors
    • Hospitality
    • Hotels
    • Housing & Social Care
    • IT
    • Legal and Compliance
    • Lifestyle
    • Marketing & Advertising
    • Media
    • News
    • Pets
    • Property
    • Real Estate
    • Research & Development
    • Retail
    • social
    • Society & Culture
    • Sports
    • sustainability
    • Technology
    • Trade
    • Transport
    • Travel
    • UK Politics
    • Vehicle
    • Weather & Climate
    • Wildlife
    • World
    Facebook X (Twitter) LinkedIn
    • Home
    • About
    • Meet the Daily News Team
    • Contact
    • Terms and Conditions
    • Privacy Policy
    © 2026 dailyNews.

    Type above and press Enter to search. Press Esc to cancel.