By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Felly ViralFelly ViralFelly Viral
Notification Show More
Font ResizerAa
  • Home
    • Home 2
    • Home 3
    • Home 4
    • Home 5
  • U.K News
    U.K News
    Politics is the art of looking for trouble, finding it everywhere, diagnosing it incorrectly and applying the wrong remedies.
    Show More
    Top News
    Latest News
  • Technology
    TechnologyShow More
    Nigeria’s regulator asks Uber about the services it did not deliver before leaving
    September 6, 2026
    Kenya bet its graduates on gig work in 2022. ChatGPT launched that year
    September 6, 2026
    The man paid to prove Fermat by hand says Claude did it in 11 days
    September 6, 2026
    Berlin refused to pay. Now it runs four task forces
    September 6, 2026
    The Fairphone 6 Plus is the midrange phone we desperately needed
    September 6, 2026
  • Posts
    • Post Layouts
    • Gallery Layouts
    • Video Layouts
    • Audio Layouts
    • Post Sidebar
    • Review
      • User Rating
    • Content Features
    • Table of Contents
  • Pages
    • Blog Index
    • Contact US
    • Search Page
    • 404 Page
    • Customize Interests
    • My Bookmarks
  • Join Us
Reading: OpenAI’s AGI number came from a harness, not the model
Share
Font ResizerAa
Felly ViralFelly Viral
  • ES Money
  • U.K News
  • The Escapist
  • Entertainment
  • Science
  • Technology
  • Insider
Search
  • Home
    • Home News
    • Home 2
    • Home 3
    • Home 4
    • Home 5
  • Categories
    • Technology
    • Entertainment
    • The Escapist
    • Insider
    • ES Money
    • U.K News
    • Science
    • Health
  • Bookmarks
    • Customize Interests
    • My Bookmarks
  • More Foxiz
    • Blog Index
    • Sitemap
Have an existing account? Sign In
Follow US
Felly Viral > Blog > Technology > OpenAI’s AGI number came from a harness, not the model
Technology

OpenAI’s AGI number came from a harness, not the model

admin
Last updated: September 6, 2026 12:23 pm
admin Published September 6, 2026
Share
SHARE
September 6, 2026 at 12:23 pmIn: Technology

A screenshot from OpenAI’s video announcing GPT-6 Astra. Credit: OpenAI OpenAI declared the AGI era on the strength of a 99.9% score. Run the same model through the benchmark’s own software, and it scores 62.7%. The gap is not a rounding error or a rival’s complaint.

It comes from the organisation that built the test. ARC Prize published both numbers on the day GPT-6 Astra launched. It printed a full table of every reasoning level it ran. The difference is the harness.

A harness is the software around a model. It sets the tools the model can reach, what it remembers between requests, and how its context gets managed. Same weights, different scaffolding, different score. What the table actually shows ARC Prize ran Astra two ways.

Its standard harness gives every model the same minimal interface and lets the model decide which notes to carry forward. OpenAI’s Provider Adapter preserves the model’s opaque reasoning state between requests and compacts longer conversations. Under the standard harness at maximum reasoning, Astra scored 62.7% and cost $26,098. Under the Provider Adapter at high reasoning, it scored 99.9% for $18,817.

The better score was also the cheaper run. One row makes the point harder than any argument. Set the reasoning effort to none inside OpenAI’s adapter and Astra still scores 96.7%. That beats the same model at maximum reasoning inside the standard harness by 34 points.

The scaffolding outperformed the reasoning dial outright. Both harnesses solved 167 game-reasoning pairs. On those, ARC Prize clocked the adapter runs at 49% fewer tokens and roughly 3.66 times faster. The number that travelled The figure that spread was 99.9% against GPT-5.6 Sol’s 7.8%.

Those are not the same test. Astra’s 99.9% came from the Provider Adapter; Sol’s 7.8% came from the standard harness. That comparison ran widely, and it ran here. TNW’s own launch coverage carried the 99.9% figure against Sol’s 7.8% without the harness caveat.

Our follow-up on the AGI claim described human parity on the same basis. The like-for-like comparison is 62.7% against 7.8%. That remains an enormous jump, and it is the one the benchmark supports. ARC Prize itself declined the conclusion OpenAI drew.

It stated plainly that it is “not claiming that it is AGI”. Co-founder Mike Knoop wrote that “we lack evidence to call this AGI yet”. Going forward the foundation will publish both harness results side by side. There is a small wrinkle inside ARC Prize too.

François Chollet gave the standard-harness figure as 66% in a post, against 62.7% in the published table. The blog is the primary record. Then the numbers moved Fortune found that the scores kept changing after publication. Emily Forlini compared archived snapshots of OpenAI’s launch post and found five metrics altered.

Astra’s hallucination rate read 4.2% in the first snapshot. By 5.20pm it read 2%. It has since gone back to 4.2%. Anthropic’s Fable 5.1 dropped nearly ten points on FrontierMath, from 87.8% to 78%, then settled at 83%.

Sol’s ExploitBench score doubled from 5.5% to 11.5%. OpenAI told Fortune it is investigating reverting that one. The 11.5% reflects a reasoning level Sol does not offer commercially. The pre-publication draft sent to media under embargo put ARC-AGI-3 at 98.6%.

The live post says 99.99%. Not every edit favoured Astra. Two Anthropic scores on HealthBench Professional went up. Other outlets reached the fine print independently, and The New Stack ran the two harness figures side by side the following day.

But the pattern runs mostly one way. OpenAI also pulled the post after publishing it and put it back up, for reasons the company said it could not disclose. Most evaluations carry noise of a few percentage points, the company told Fortune, depending on checkpoint, scaffold and evaluation run. Two Stanford researchers have a word for it Anka Reuel and Mike Hardy call the practice benchmaxxing, meaning re-running evaluations under different conditions until the number improves.

Both work at Stanford, at the Intelligent Systems Laboratory and the Trustworthy AI Lab. They also went looking in the system card, which is where the method should be documented. On the internal hallucination benchmark they found “barely any details about the evaluation”, adding that it “doesn’t even include the number of test items”. Not everyone reads it as gaming.

Vincent Sunn Chen of Snorkel AI told Fortune that scores routinely shift in the final hours before a launch. Checkpoint, configuration, harness and grading are all still moving. His suggestion is a norm requiring companies to say what changed when they revise a published figure. There is precedent for the harsher reading.

Critics accused Meta in 2025 of publishing Llama 4 results from an internal build rather than the public one. Yann LeCun later said the company had fudged them. Where Astra actually sits Artificial Analysis ran its own tests and got a duller picture. On its Coding Agent Index, Astra scores 67 in Codex, level with Claude Opus 5 and Fable 5.

Fable 5.1 leads at 70. On its Intelligence Index, Astra scores 61, the same as the model it replaces, five points behind Fable 5.1 and behind Meta’s Muse Spark 1.3. It costs $10 per million input tokens and $50 per output, two and a half times Sol’s price, which works out 75% more expensive per task at maximum effort. The gains are real but narrow.

Hallucination on its knowledge benchmark fell from 92% to 51%, and long-horizon knowledge work rose around 80 Elo. An economically-weighted task benchmark fell by roughly the same amount, with further regressions on banking support, scientific Python and long-context reasoning. Artificial Analysis has since rebuilt the index. Version 4.2 arrived the following day with harder tasks and more private test sets, which the firm says exist to prevent gaming.

The harness is the product now We have run into this word twice already. A cheap harness collapsed Booz Allen’s own league table of AI threat. CrowdStrike chose the same word for what it wraps around OpenAI’s cyber model. Anthropic, Google and Microsoft all sell harnesses as products with their own pricing.

Nvidia built one that took Claude Opus 5 from 30.2% on ARC-AGI-3 to clearing every level. OpenAI is not hiding anything here. The adapter uses documented API features any developer can call. But the thing that scored 99.9% is an assembled system, and the thing on sale is a model.

That distinction is the whole story, and it also runs through the monitorability argument. What a model can do and what anyone can independently verify it doing are drifting apart, and the benchmark is where you can see the gap.

You Might Also Like

Nearly impossible? How Fairphone built the ethical, repairable Fairphone Gen 6+.

Once popular for attacking AI, ASCII smuggling is embraced by spammers

Nigeria’s regulator asks Uber about the services it did not deliver before leaving

The Fairphone 6 Plus is the midrange phone we desperately needed

Tesla’s Cybercab has been deployed, and it’s already under investigation

Share This Article
Facebook Twitter Email Print
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

What's Hot

Badenoch accuses Burnham of kicking defence spending into the long grass

Zack Polanski says he'll run for former PM Starmer's seat in by-election

Hollie Doyle blog: How my husband Tom can clinch Sprint Cup title

O'Neill satisfied with Celtic signings | McInnes has Rangers squad he craved

'It's going to be tight' – Hamilton eyes 'dream' Monza pole

Digested week: Tom Tugendhat is a serious politician. What’s he doing in a Reform-lite Tory frontbench? | John Crace

Categories

Business

11 Articles

Celebrity

Fashion

Politics

32 Articles

Travel

- Advertisement -
Ad image

Categories

  • ES Money
  • U.K News
  • The Escapist
  • Insider
  • Science
  • Technology
  • LifeStyle
  • Marketing

About US

We influence 20 million users and is the number one business and technology news network on the planet.

Subscribe US

Subscribe to our newsletter to get our newest articles instantly!

© Foxiz News Network. Ruby Design Company. All Rights Reserved.

Powered by
►
Necessary cookies enable essential site features like secure log-ins and consent preference adjustments. They do not store personal data.
None
►
Functional cookies support features like content sharing on social media, collecting feedback, and enabling third-party tools.
None
►
Analytical cookies track visitor interactions, providing insights on metrics like visitor count, bounce rate, and traffic sources.
None
►
Advertisement cookies deliver personalized ads based on your previous visits and analyze the effectiveness of ad campaigns.
None
►
Unclassified cookies are cookies that we are in the process of classifying, together with the providers of individual cookies.
None
Powered by
Welcome Back!

Sign in to your account