OpenAI announced its new Astra model on September 3, but the rollout of the accompanying blog post did not go smoothly. The post was originally scheduled for 2 p.m. Eastern, was pulled shortly after it went live, and reappeared around 4 p.m. with different evaluation numbers than the version that briefly existed before it. A promotional tweet sent at 3:32 p.m. returned error messages, and CEO Sam Altman posted at 3:50 p.m. that "we hit a little snag getting the blog post deployed." OpenAI initially attributed the delay to a bug in its content management system, then to an internet outage, without offering a fuller explanation of what actually happened.
What has drawn more attention than the outage itself is how many of the benchmark figures changed between versions. Astra's reported hallucination rate was listed at 4.2 percent in the original post, dropped to 2 percent in the revised version, then returned to 4.2 percent in the version now live. A separate OpenAI model, GPT-5.6 Sol, followed the identical pattern on the same metric, moving from 12.2 percent to 9.4 percent and back to 12.2 percent.
Math scores on FrontierMath's hardest tier told a similar story. Astra's own result held steady at 97.6 percent across every version, but the figures OpenAI cited for competing models moved considerably: Anthropic's Fable 5.1 was listed at 87.8 percent, then 78 percent, then 83 percent across the three versions of the post, while Sol's own FrontierMath score went from 83 percent to 80.5 percent and back to 83 percent. On ExploitBench, a cybersecurity benchmark, Sol's score jumped from 5.5 percent to 11.5 percent, a change OpenAI is reportedly now reviewing whether to revert. A pre-publication draft had also listed Astra's score on ARC-AGI-3 at 98.6 percent; the published post put it at 99.99 percent.
OpenAI has not given a detailed public account of why so many numbers shifted across three versions of the same announcement. A company spokesperson said only that "we care deeply about getting evaluations right" and that fixes were made "to ensure the numbers represent our best estimate."
That explanation has not satisfied everyone watching the AI benchmarking world closely. Stanford researchers Anka Reuel and Mike Hardy have raised the possibility that OpenAI re-ran its tests repeatedly in search of better scores before publishing, a practice some in the field have started calling "benchmaxxing." Vincent Sunn Chen of Snorkel AI took a more measured view, noting that scores commonly shift in the days before a model launch as testing continues, but argued that companies owe the public a clear explanation of exactly what changed and why once numbers are made public.
The episode revives a debate that flared up in 2025, when Meta's then-chief AI scientist Yann LeCun acknowledged the company had "fudged" benchmark results for its Llama 4 model. As AI companies compete for enterprise customers, and in some cases prepare for eventual public offerings, published benchmark numbers have become a kind of currency in their own right, one that a moving target does little to strengthen.

