OpenAI's GPT-6 Astra on ARC-AGI-3 (arcprize.org)

185 points by vignesh_warar 9 hours ago

at1as 5 hours ago

I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken.

From https://epoch.ai/latest/announcing-frontiermath-erdos

> Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours

> Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.

Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.

zone411 2 hours ago

I don't:

"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.

Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.

red75prime 2 hours ago

A very long tail of problems that weren't solved by humans? Sure.

It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.

at1as an hour ago

What does this mean?

There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems

Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof).

I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve

red75prime an hour ago

malfist 8 hours ago

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

matherial 8 hours ago

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.

I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

howunfortunate 3 hours ago

IQ tests are incredibly good at what they're designed for, which is discriminating relatively higher intelligence humans from lower intelligence humans. Also discriminating within a single human - they are routinely and reliably used to track cognitive decline.

For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).

They were never designed for machines or non-human animals.

Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.

pavlov 7 hours ago

When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.

It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

phainopepla2 6 hours ago

guelo 7 hours ago

sfblah 7 hours ago

tintor 7 hours ago

Pure software benchmarks might be getting saturated, but physical ones aren't.

Let LLM control a physical robot to perform tasks that average human can do.

quotemstr 4 hours ago

skybrian 7 hours ago

Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.

idiotsecant 4 hours ago

ranyume 7 hours ago

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

guelo 7 hours ago

jawiggins 8 hours ago

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

eli 8 hours ago

Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.

ranyume 7 hours ago

Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.

malfist 8 hours ago

I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

hyperhello 8 hours ago

whattheheckheck 7 hours ago

paimapi 8 hours ago

I'm also unclear as to how basic inferential logic puzzles spells out intelligence

I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

mdp2021 7 hours ago

> basic inferential logic puzzles spells out

It spells out a form of intelligence - some can and some cannot.

Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.

paimapi 4 hours ago

jrflo 8 hours ago

You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

https://arcprize.org/

rcoveson 8 hours ago

No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.

malfist 8 hours ago

You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

mdp2021 7 hours ago

dist-epoch 8 hours ago

1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.

GaggiX 8 hours ago

If you have never seen the game before probably.

Betelbuddy 9 hours ago

"For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

Well I dont know about all of you, but I am celebrating meat based humans...

LPisGood 8 hours ago

I think raw brain energy is not a fair comparison. Humans are not willing and able to serve requests at identical competence all hours of the day. You have to invest considerable resources to get a person to even do so for part of the day.

paxys 8 hours ago

Why are you making the assumption that a person's time is worthless? I'd argue that it is the single most valuable resource we all have.

modeless 5 hours ago

$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.

fastball 6 hours ago

Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?

an0malous 6 hours ago

Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.

an0malous 5 hours ago

Actually I can answer my own question: we know that they have had previous access to the tests because they’ve run older models against the same benchmark.

I wouldn’t put it past a company like OpenAI with a long history of lying and being deceptive to record the tests and benchmaxx ARC. They have trillions of dollars of incentive to cheat any way they can.

6thbit 7 hours ago

The instant/no reasoning performed extremely well

    none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.

NitpickLawyer 7 hours ago

Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).

dwohnitmok 8 hours ago

> Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

mikert89 8 hours ago

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

tedsanders 7 hours ago

Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

ranyume 7 hours ago

Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.

mikert89 7 hours ago

these just need more compute:

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

but we both know these examples go against the spirit of my point

tedsanders 7 hours ago

jdthedisciple 6 hours ago

Yes, but not necessarily under tight budget constraints.

mikert89 2 hours ago

theres no budget constraints for AGI

x3haloed 8 hours ago

Yup. Only subjective taste remains.

GPerson 8 hours ago

Nope that will be commodified in short order.

fxd 6 hours ago

“AGI” never made sense to me. It’s a purely marketing term right?

I’ve ignored it thinking it would go away, but it keeps coming up.

I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery.

Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get it.

But you’d be no nearer to solving consciousness.

Given this thought trajectory - what is AGI supposed to be?

layer8 6 hours ago

Not sure why you are bringing up consciousness, that’s largely orthogonal to intelligence. AGI is usually taken to mean the capability to match or surpass human intelligence across all conceivable cognitive tasks, as opposed to being limited to certain kinds of tasks, or to not matching the general level of human intelligence in some respect.

Intelligence, and hence AGI, doesn’t require consciousness or emotions or sentience.

p1esk 6 hours ago

capability to match or surpass human intelligence across all conceivable cognitive tasks

What human intelligence do you mean? Genius? Professional? Educated? Random person? “Dumb” person?

drdeca 5 hours ago

layer8 6 hours ago

meander_water 6 hours ago

The OpenAI charter defines it as:

"highly autonomous systems that outperform humans at most economically valuable work"

https://time.com/article/2026/08/26/openai-sam-altman-interv...

fxd 3 hours ago

So vague it’s useless.

We cannot in a declarative sense define what is economically valuable work even now let alone into the future.

People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”.

Of course an AI will accept $0/hr, it doesn’t mean it does the job.

Anyone who could accurately define the value of work would be wildly successful without having to try.

That is not a useful definition for me unfortunately.

jryle70 2 hours ago

mdp2021 4 hours ago

> Knowledge and thus intelligence

How can you conflate the two.

> solving consciousness

We are very much not interested in that. We just need a proper problem solver.

fxd 4 hours ago

Right, you are interested in “AGI” and presuming none of that requires consciousness right?

For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain

submain 2 hours ago

Unproven, but it could very well be some problems require consciousness to be solved.

eagerpace 6 hours ago

I like recursive self improvement instead. It seems like something that is actually quantifiable and kinda “the point” of why consciousness is important to humans.

fxd 6 hours ago

So basically, being able to set it free on some long running goal and it sort of “lives” and autonomously does its own tasks?

I wonder at what point consciousness is necessary… that is, if you can have anything like that without it.

To the point that solving consciousness (and combining it with intelligence) is what gives you the autonomous, recursive, self-improving thing otherwise it can only drive in the dark and make big mistakes.

To your point I think - it’s why we don’t see too many non-conscious advanced biology (it rarely survives against those with it).

scotty79 an hour ago

How good are LLMs at doing Mensa tests?

brokensegue an hour ago

IQ tests? Very good. But most are in the dataset so it's not very meaningful

piloto_ciego 9 hours ago

99.9% with the right harness? Ok, we're at AGI then.

Prediction:

We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing about if the singularity is here and if we're at AGI for the rest of my life most likely.

WASDx 7 hours ago

They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".

emp17344 8 hours ago

Then why is unemployment around 4%? You believe we have AGI and yet it can’t do anyone’s job?

piloto_ciego 8 hours ago

Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?

raspasov 7 hours ago

jhonof 9 hours ago

NitpickLawyer 7 hours ago

AFAICT nvda's result is on the 25 open problems, while this submission is on the "semi-private" set, ran by the arc people themselves.

piloto_ciego 9 hours ago

I rest my case.

dgellow 8 hours ago

The goal moving is by design, that’s why they use something as ill defined as AGI

slopinthebag 9 hours ago

That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.

piloto_ciego 8 hours ago

I agree, like the average human isn't generally intelligent.

IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

baal80spam 9 hours ago

> We will now see the goalposts moved

It's already happening :)

piloto_ciego 8 hours ago

Hilariously it is, I'm just reading more on this!

yomismoaqui 6 hours ago

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI?

Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

hypfer 8 hours ago

What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?

petu 8 hours ago

OpenAI provides API key with ~unlimited use?

Frost1x 8 hours ago

So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

Alignment++

manquer 4 hours ago

yusufozkan 9 hours ago

what the hell is that score/cost curve lol

minimaxir 9 hours ago

DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.

Frost1x 8 hours ago

It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

Phemist 8 hours ago

What is the intuition. Higher quality turns due to more reasoning results in significantly fewer turns taken?

tedsanders 7 hours ago

minimaxir 8 hours ago

bigbuppo 7 hours ago

Wake me when it's going to spontaneously fix my leaky faucet because if it doesn't do it nobody else will. Until it has that capability I don't really care.