The unveiling of OpenAI’s GPT-6 Astra model was not without its peculiar circumstances, marked by a blog post that struggled to go live and, more notably, a series of evolving performance metrics that have raised questions within the AI community. Originally slated for a 2 p.m. ET release on September 3, the announcement blog encountered technical snags, delaying its widespread visibility online for nearly two hours. Even after OpenAI CEO Sam Altman shared a link, many users, including journalists at Fortune, reported error messages. When the post finally became consistently accessible, it contained evaluation figures that differed from earlier versions, and these numbers have continued to fluctuate even after publication.
A striking example of these alterations concerns Astra’s reported hallucination rate. Early archived snapshots of the blog post, taken around 2:23 p.m. ET, showed Astra’s hallucination rate at 4.2%. This figure persisted through several subsequent snapshots, yet by 5:20 p.m., after the blog was widely viewable, the rate was halved to 2%. Curiously, scores for Astra’s predecessor, GPT-5.6 Sol, also saw a reduction, dropping from 12.2% to 9.4% in the same period. More recently, these hallucination rates for both models have reverted to their original, higher percentages of 4.2% and 12.2% respectively, according to the source material. Such dynamic adjustments invite scrutiny into the precise methodology and stability of these critical performance indicators.
Another area where numbers shifted was in mathematics capabilities, a strength OpenAI specifically highlighted for Astra. While Astra’s score on the FrontierMath Tier 4 (v2) evaluation remained constant at 97.6%, scores for GPT-5.6 Sol and Anthropic’s Fable 5.1 model were briefly altered. Fable 5.1’s score, initially 87.8% in the earliest snapshot, dipped to 78% by late afternoon before climbing back to 83%. Similarly, GPT-5.6 Sol’s math score transitioned from 83% down to 80.5% and then returned to 83%. These changes, even if temporary, created an impression of Astra performing significantly better against its rivals at certain points, underscoring the delicate nature of public-facing benchmarks.
The adjustments extend beyond just these metrics. An embargoed draft provided to media outlets prior to the official launch showed Astra’s ARC-AGI-3 evaluation score at 98.6%, a figure that now stands at 99.99% on the live blog. OpenAI attributed such changes between draft and final versions to normal verification processes. A spokesperson noted that the Arc Prize Foundation’s independent assessment found Astra performing at 99.9% when given a powerful harness, though it achieved 63% with the standard harness. These nuances, involving factors like the “harness” or “reasoning level,” highlight the complexity in standardizing AI model evaluations.
This phenomenon of shifting metrics is not entirely new to the AI industry, which grapples with intense competition and the challenge of accurately measuring large language model performance. Researchers Anka Reuel and Mike Hardy from Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab pointed to “benchmaxxing” as a known practice, where evaluations are re-run under varied conditions to maximize scores. They also noted that the GPT-6 Astra system card offered limited detail on certain evaluations, such as the internal hallucination benchmark, raising concerns about transparency.
An OpenAI spokesperson stated that the company “cares deeply about getting evaluations right” and that “fixes were made to ensure the numbers represent our best estimate of available model performance.” They explained that most evaluations have inherent “noise within a few percentage points.” Vincent Sunn Chen, an AI engineer specializing in benchmarks, echoed that shifts in scores are not uncommon in the final hours before a model launch, often due to evolving configurations and measurement setups. However, the continuous post-launch changes, particularly those that initially favored Astra before reverting, present a complicated picture for an industry where benchmark leadership can heavily influence customer perception, investor confidence, and even recruitment of top talent. The ongoing fluidity of these reported figures could ultimately muddy the waters for those seeking clear, consistent data on the true capabilities of cutting-edge AI.
