# OpenAI Revises GPT-6 Astra Benchmark Scores After Launch

> Updates to evaluation metrics for the frontier model reveal a substantial gap between reported figures and independent tests, highlighting issues with evaluation consistency amid rising competition.

*Published 2026-09-06 · By Marcus Vance*

GPT-6 Astra is OpenAI's frontier model whose benchmark scores underwent multiple post-launch revisions on evaluations such as ARC-AGI-3.

OpenAI updated several benchmark scores for GPT-6 Astra in its launch blog after the initial publication on September 3, 2026. The adjustments included improvements to the model's reported performance on certain metrics and temporary reductions to scores associated with competing models from Anthropic. These changes occurred amid ongoing competition in the development of advanced AI systems. The company provided an embargoed draft to media outlets prior to the public release. That draft contained different figures for key evaluations compared to the final published version. Similar patterns of fluctuation appeared in records for the predecessor model GPT-5.6 Sol. The revisions focused on metrics that assess reasoning capabilities and error rates.

## What specific benchmark changes occurred in the GPT-6 Astra launch materials?

The public launch blog listed GPT-6 Astra at 99.9 percent on ARC-AGI-3. The embargoed draft had listed the same model at 98.6 percent on that benchmark. OpenAI also adjusted hallucination rates across archival snapshots. Early versions showed 4.2 percent while a later snapshot recorded 2 percent before the figure returned to 4.2 percent. The company stated that most evaluations carry noise within a few percentage points depending on the exact checkpoint, scaffold, and evaluation run. These adjustments were presented as corrections to better represent available model performance for user comparisons. The launch materials also noted that Astra surpassed the human action-efficiency baseline on 96 percent of levels in the ARC-AGI-3 benchmark.

The revisions extended to other undisclosed metrics in the blog post. Some updates favored GPT-6 Astra while others lowered figures for Anthropic models including Claude Fable 5.1 and Claude Opus 5. The timing of these changes followed a rare delay in the publication of the announcement. Independent observers noted the shifts through comparisons of the draft materials and the final blog. The process illustrates how reported numbers can evolve between internal sharing and public disclosure. Such evolution affects how the model is positioned relative to other frontier systems released around the same period.

## How do the official scores compare against independent ARC Prize results?

ARC Prize reported GPT-6 Astra at 62.7 percent on ARC-AGI-3 when using the standard harness that carried an estimated cost of 26,000 dollars. The same model reached 99.9 percent when tested with the provider adapter harness. This difference accounts for the approximately 37-point gap referenced in coverage of the revisions. The standard harness applies uniform testing conditions across models without model-specific adaptations. The provider adapter harness incorporates optimizations supplied by the model developer. These methodological differences lead to divergent outcomes on the same underlying benchmark tasks. The ARC Prize findings were published separately from the OpenAI materials and provide an external reference point.

The gap between the two testing approaches raises questions about which harness best reflects real-world deployment conditions. Standard harness results may offer more comparable data across different developers. Adapter-based results can highlight maximum potential under tailored conditions. The 62.7 percent figure from the standard harness contrasts sharply with the 99.9 percent figure from the OpenAI blog. This contrast appears in the ARC Prize analysis of the model. The discrepancy underscores the need for clear disclosure of testing conditions when benchmark numbers are released to the public.

GPT-6 Astra evaluation metrics reported across OpenAI and ARC Prize sourcesBenchmarkOpenAI Public BlogEmbargoed DraftARC Prize Standard HarnessARC-AGI-399.9%98.6%62.7%Hallucination Rate4.2%Varied to 2% then 4.2%Not specified

## What factors contribute to the observed variations in reported scores?

OpenAI identified harness selection, reasoning level settings, and evaluation run specifics as sources of variation. The company noted that small differences in these parameters can shift results by several percentage points. The provider adapter harness allows integration of model-specific tools that are not present in the standard harness used by ARC Prize. This distinction explains much of the 37-point difference on ARC-AGI-3. Similar patterns of change appeared in hallucination metrics across multiple snapshots of the same model. The predecessor GPT-5.6 Sol exhibited comparable fluctuations in its reported hallucination rates. These observations suggest that benchmark reporting involves choices that influence final published numbers.

The company emphasized that its launch blog figures represent the best estimate of model performance after internal fixes. The spokesperson statement addressed concerns about consistency by pointing to inherent noise in evaluation processes. The statement appeared in coverage of the post-launch adjustments. The approach allows users to make comparisons based on the adjusted numbers. However, the availability of independent results using different harnesses provides additional context for interpreting those numbers. The combination of internal revisions and external tests creates a more complete picture of performance claims.

- Media received an embargoed draft containing the initial 98.6 percent ARC-AGI-3 score.
- OpenAI published the launch blog on September 3, 2026, with the ARC-AGI-3 score updated to 99.9 percent.
- ARC Prize released independent test results showing 62.7 percent under standard harness conditions.
- Archival snapshots revealed hallucination rate changes from 4.2 percent to 2 percent and back to 4.2 percent.
- Temporary downward adjustments appeared for select Anthropic model scores in the updated blog.

## What market and stakeholder implications arise from the benchmark revisions?

The revisions affect how investors, developers, and enterprise users assess the relative capabilities of frontier models. A reported score of 99.9 percent on ARC-AGI-3 positions GPT-6 Astra as reaching human parity on that benchmark according to OpenAI launch materials. The independent 62.7 percent result under standard conditions presents a different view of the same model. Stakeholders must therefore weigh which testing methodology aligns with their intended use cases. The temporary adjustments to Anthropic model scores add another layer of complexity to direct comparisons between competing systems. Such adjustments can influence market perceptions during periods of rapid model releases.

Enterprise customers evaluating models for deployment consider hallucination rates alongside reasoning benchmarks. The documented fluctuations in that metric for both GPT-6 Astra and GPT-5.6 Sol indicate that even internal records can vary over short periods. This variability may prompt organizations to request raw evaluation data rather than relying solely on published summaries. The situation also draws attention to the role of third-party testers like ARC Prize in providing standardized reference points. Overall, the episode illustrates how benchmark presentation can shape competitive positioning in the frontier model sector.

> We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.OpenAI spokesperson

## What reactions have emerged from the AI research community?

The OpenAI spokesperson response framed the changes as necessary corrections to improve the accuracy of published numbers. The statement stressed the importance of enabling meaningful comparisons among available models. Coverage in Fortune documented the sequence of draft and public versions along with the archival snapshot changes. The ARC Prize report supplied the contrasting independent data that highlighted the harness effect. These two sources together provide the primary documented accounts of the events. No additional external expert commentary appears in the available materials.

## What developments may follow for frontier model evaluations?

The documented differences between harness types suggest that future reporting may include explicit details on testing conditions. Clearer separation of standard harness results from adapted results could improve transparency. The 37-point gap on ARC-AGI-3 demonstrates how methodological choices affect final figures. Stakeholders may increasingly seek multiple evaluation perspectives before forming conclusions about model capabilities. The pattern of post-launch adjustments also points to the value of maintaining versioned records of benchmark claims. Such practices could help track how reported performance evolves after initial announcements.

The situation with GPT-6 Astra and its predecessor GPT-5.6 Sol shows that even within one organization, metrics can shift across snapshots. This observation applies to both reasoning benchmarks and hallucination rates. Continued scrutiny from independent labs may encourage more standardized approaches across the industry. The involvement of ARC Prize in providing harness-specific results offers one model for external validation. Over time, these dynamics could lead to greater emphasis on reproducibility in frontier model evaluations. The current case supplies a concrete example of the challenges involved in maintaining consistent reporting standards.

## Sources

1. [Primary OpenAI launch announcement detailing benchmarks including 99.9% on ARC-AGI-3 and other scores.](https://openai.com/index/gpt-6-astra/)
2. [Independent lab report from ARC Prize on Astra's performance under different harnesses: 62.7% standard vs 99.9% adapter.](https://arcprize.org/blog/astra)
3. [Details on post-launch changes, embargoed draft scores, archival snapshots, and OpenAI spokesperson response.](https://fortune.com/2026/09/04/openai-quietly-boosts-some-of-astras-evaluation-metrics-amid-rare-delay-in-publication-of-the-modeblog-post-announcement/)

---
Source: https://aiintelreport.com/frontier-models/openai-revises-gpt-6-astra-benchmarks
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
