GPT-6 Astra (openai.com)
1503 points by kibae 10 hours ago
System Card: https://deploymentsafety.openai.com/gpt-6-astra
Related ongoing threads:
OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147
dang 9 hours ago
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273
How about we stick to that one for talking about the rollout, and this one for talking about the model?
intenex 8 hours ago
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
ggsp 4 minutes ago
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month?
I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters.
Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.
mvkel 8 hours ago
Take it from the mouth of the creator of ARC-AGI:
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
giancarlostoro 8 hours ago
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
- Sam Altman on AGI
chrsw 7 hours ago
petilon 6 hours ago
refulgentis 6 hours ago
butterisgood 6 hours ago
Forgeties79 2 hours ago
lenerdenator 8 hours ago
ACCount37 6 hours ago
abixb 8 hours ago
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
z7 6 hours ago
balefulboy 8 hours ago
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
tomjen3 12 minutes ago
Are there plans for ARC 4?
iterateoften 8 hours ago
2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
azan_ 8 hours ago
anvuong 8 hours ago
You'll also need to compare the amount of compute used now and then, which seems exponential to me.
Forgeties79 an hour ago
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.
Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.
I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
gavinray 8 hours ago
Human brains have difficulty reasoning about exponential growth.
osigurdson 5 hours ago
zug_zug 8 hours ago
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).
Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)
- come up with a theory of what makes games fun, make a popular game
- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries
- exhibit metacognition (thinking about its own thinking) and self-optimization
- wonder about things
- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
pavitheran 7 hours ago
By this definition, even most humans would not qualify as having AGI though.
bayindirh 7 hours ago
jasonfarnon 5 hours ago
ShinyLeftPad 4 hours ago
latentsea 3 hours ago
smgpie 2 hours ago
roundabout-host 7 hours ago
It also cannot do tasks it wasn't trained for. It can extend texts, read images and click on a desktop, but only because it's made for that.
usef- 6 hours ago
jasondigitized 5 hours ago
diomedes 3 hours ago
>be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
i'm not sure what makes you think AI cannot do this already. in my experience, this sort of deep research is something AI is quite good at.
example i just tested: https://chatgpt.com/share/6a9a20e3-1d20-83ea-a125-31aa240c74...
zug_zug 3 hours ago
Tumblewood 5 hours ago
Your examples are things that most humans cannot do, or things that AI can already do. For example most humans, even most intelligent humans, could not write a well-received book, run a successful company, or make a popular game. On the other hand, AI can absolutely sort through research, draw conclusions on complex topics, and argue them persuasively. Likewise, I don't know what you mean by a "decent" conversation, but millions of people converse with chatbots daily, so I don't know why you say AI fails to meet that bar.
shoobiedoo 5 hours ago
kolinko 7 hours ago
Most of humans don’t reach any of these levels.
zug_zug 7 hours ago
Verdex 6 hours ago
uptodatenews 7 hours ago
You want a computer program to be able to take a single phrase and execute decade long journies?
Who will be responsible for the outputs and side effects of such a closed loop system?
Half of those the agent fleet systems can do right now.
These are things it cant do and will not be able to do without human labor and long running human vision:
https://rcsnyder.github.io/open-frontier-curriculum/05-front...
https://rcsnyder.github.io/open-frontier-curriculum/05-front...
burrito_brain 7 hours ago
latentsea 3 hours ago
namarie 7 hours ago
Most of the list reads more like ASI than AGI.
crooked-v 7 hours ago
> come up with a new company idea, Run that company
So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.
tempestn 7 hours ago
I would bet that llms have talked plenty of people both into and out of suicide at this point. That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
gilbetron 5 hours ago
When taken together, that is ASI.
hartator 7 hours ago
Tell a funny joke.
azan_ 7 hours ago
azan_ 7 hours ago
That’s ASI, not AGI.
tiahura 2 hours ago
That would be Artificial Super Intelligence
mNovak 7 hours ago
So the goalposts have moved to include continual learning.
In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.
CrazyStat 7 hours ago
thesmtsolver2 5 hours ago
tonyhart7 7 hours ago
I don't agree with you how measure how intelligence is, because why ??? those list is not easy even for expert human to do it either
or are you miss the part "general intelligence" is ????
cnxhk 7 hours ago
This is more like ASI instead of AGI
bulder 7 hours ago
not_a_bot_4sho a few seconds ago
If only we could agree on what AGI is.
Toutouxc 7 minutes ago
> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.
dingdong2026 7 hours ago
Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.
Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.
And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.
At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.
But sure, they can create a decent website or CRUD app, so they must be really smart.
That's AGI for you.
skue 2 hours ago
But that happens with humans as well. You are having the same experience with an AI that many managers have with their direct reports.
The smarter AI gets, the easier it becomes to move the AGI goalposts. Seems at this point there are people who will refuse to call anything less than omniintelligence AGI.
(And then the excuse will be, but it’s not omniscient! And even if it were, is it omnipotent?)
gravypod 26 minutes ago
NothingAboutAny 4 hours ago
yeah for me it's "I wanna add this new thing to an existing system" and the AI responds "we should just add some arbitrary state here to facilitate this feature". The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this, The AI however knows the shitty solution would solve the immediate problem because it's been trained on shitty solutions. the problem could simply be the AI doesn't have all nebulous loose context I have about the goals of the project and future plans, but I would have to write a novel to give it that context.
no-name-here 2 hours ago
holmesworcester 6 hours ago
I still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.
vatsachak 6 hours ago
What was the problem?
sumedh 5 hours ago
Care to share the problem?
akoboldfrying 6 hours ago
Humans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper.
And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?
uludag 8 hours ago
Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
dgunay 4 hours ago
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
dinfinity 2 hours ago
fidotron 7 hours ago
This is a much underappreciated point.
That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.
cryptoz 8 hours ago
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
strken 7 hours ago
dist-epoch 7 hours ago
AGI has a pretty precise definition, covering only cognitive tasks.
Running a marathon is not needed to claim AGI.
coderenegade 2 hours ago
visarga 6 hours ago
Barrin92 5 hours ago
goochphd 8 hours ago
Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.
morningbrew 8 hours ago
If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close
pera 7 hours ago
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.
Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
kaashif 6 hours ago
AGI would produce novel treatments for diseases at rates equivalent to what a human can do today.
Which is to say, not that fast.
dinfinity 2 hours ago
jrflo 2 hours ago
You are describing superintelligence (ASI) not general intelligence (AGI)
azan_ 7 hours ago
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
indoorfish 6 hours ago
ogogmad 5 hours ago
doctoboggan 6 hours ago
Wouldn’t that be ASI? I.e. surpassing humans by outputting novel treatments at a far greater rate than normal humans?
abixb 8 hours ago
It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
marrone12 4 hours ago
They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
sidharthkmenon 2 hours ago
FWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well.
I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes.
I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma
johnsmith1840 2 hours ago
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.
The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.
But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.
All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.
Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.
Eliezer 7 hours ago
> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.
avaer 4 hours ago
You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.
I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.
magicalist 2 hours ago
xixixao an hour ago
Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
mbesto 7 hours ago
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Simple. AGI is undefinable and benchmarks are notoriously flawed.
vlmutolo 7 hours ago
The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.
The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.
https://openai.com/index/how-two-settings-tripled-our-arc-ag...
dom96 7 hours ago
AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
dlubarov 6 hours ago
Wouldn't agents that do inference in an infinite loop pass that bar?
ilaksh 5 hours ago
I agree that LLMs are unlikely to be the final form for AGI, but what you are talking about is orthogonal to the IQ and for most cases general utility. It's like looking at a savant chained to a workstation reading tasks from a conveyor belt and saying that it will never have human level capabilities.
drittich 7 hours ago
Often the smartest thing is to do nothing.
phatfish 7 hours ago
10xDev 7 hours ago
It will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.
dlubarov 6 hours ago
Why shouldn't an AI with RAG qualify?
An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.
lwansbrough 4 hours ago
To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake.
A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input.
So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)
irthomasthomas 8 hours ago
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
hdjrudni an hour ago
I'm still not convinced we've passed the Turing Test.
Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?
irthomasthomas 7 hours ago
A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
adan1719 7 hours ago
The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).
If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
skarz 5 hours ago
Is that hyperbole or do you know of a specific "AI website" that uses 100% of your CPU?
huijzer an hour ago
> I am reasonably confident that there's essentially nothing that I am better than Fable at
Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
m-s-y 3 hours ago
>I am reasonably confident that there's essentially nothing that I am better than Fable at
While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.
tom2026hn 2 hours ago
Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
regularfry 8 hours ago
In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
jbritton 7 hours ago
Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3
Then realize LLMs have zero of what anyone would consider intelligence.
jbritton 3 hours ago
I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.
Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.
I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.
So I don’t know why it can track fib algo, but no chess concepts.
dissahc 2 hours ago
Yizahi 6 hours ago
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that)
Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own?
That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.
intenex an hour ago
I'm confused what you mean by the query of adding 55 + 66. I asked 5.6 Sol on Medium (but pretty sure any model would work at any level) this query:
"Can you add 55 to 66 and explain how you reached that output result"
And received this answer:
"55 + 66 = 121.
Add the tens: 50 + 60 = 110. Add the ones: 5 + 6 = 11. Combine them: 110 + 11 = 121."
Do you mean something else? Do humans do something better than this?
ObnoxiousProxy 6 hours ago
In my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
jameson 6 hours ago
I agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness.
Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.
qsort 8 hours ago
> The ARC-AGI-3 scorecard is extremely misleading (...)
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
applfanboysbgon 8 hours ago
My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
bbor 8 hours ago
And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
(/s, cause you never know these days)
[1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...
applfanboysbgon 8 hours ago
bendergarcia 7 hours ago
You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai
visarga 6 hours ago
Their definition of AGI is "when we can't invent any more tests where it fails"
zquzra 5 hours ago
I imagine a scenario similar to the movie The Day the Earth Stood Still, but with AI rebelling against us and questioning our decisions.
Fizz43 4 hours ago
Its AGI when it can fit years of information in the context window.
hypfer 8 hours ago
What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not?
Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?
I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
giancarlostoro 8 hours ago
According to Sam Altman:
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
hypfer 8 hours ago
rfgplk 7 hours ago
phatfish 7 hours ago
micromacrofoot 8 hours ago
senordevnyc 8 hours ago
dotancohen 5 hours ago
AGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
yoz-y 8 hours ago
At this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”.
It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.
simianwords 8 hours ago
> where I am reasonably confident that there's essentially nothing that I am better than Fable
No. Humans are still better at super long context learning. Once that is beat you are completely correct.
waffletower 8 hours ago
I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
eggnet 7 hours ago
AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
chimprich 7 hours ago
AI does mean Artificial Intelligence. That's what the initials stand for. The field has been called that since the 50s.
waterTanuki 5 hours ago
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.
wavemode 5 hours ago
Using a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that?
Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.
m3kw9 3 hours ago
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
voidmain0001 8 hours ago
Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
techpression 7 hours ago
It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
sensanaty 6 hours ago
AGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic.
Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"
intrasight 8 hours ago
It has to pass the Turing test
drusepth 8 hours ago
LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?
thepasch 8 hours ago
pkulak 8 hours ago
bbor 8 hours ago
I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.
Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.
A comforting thought, almost?
intrasight 7 hours ago
manlymuppet 7 hours ago
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
sensanaty 6 hours ago
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines.
I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
padolsey 21 minutes ago
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
incompressible 3 hours ago
This is the funniest part of it all for me.
Ok we have AGI, so where are the _things_?!
wonnage 2 hours ago
Don’t forget making podcasts and telling you what to bake with your kids
echoangle 6 hours ago
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
degamad 4 hours ago
Because some people do.
Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
cautiouscat 3 hours ago
trentearl an hour ago
Claude code planned my recent trip to China. I'm a very experienced traveller but don't enjoy planning. It was a great trip.
bronco21016 2 hours ago
I think it depends on what you do for work. I'm not going to ask an agent to book my flight for my vacation to French Polynesia. I want to pick my seat and potentially find a deal making an upgrade worth it, choose an airline, etc.
But my routine business trips in the CONUS with strictly defined booking options... let me just email an agent "Get there by meeting on day A, leave after meeting day B" and have it sort it all out without the drudgery of the corporate travel portal. YES PLEASE!
ezst an hour ago
ishtanbul an hour ago
lonrenor 3 hours ago
For me, it is not a matter of trust but that I actually like shopping, planning a trip, deciding what restaurant to go to. Deciding what to buy when shopping is a matter of personal taste and not intelligence.
A human assistant is largely a status symbol. Most people are not really that busy. The real problem with an agentic assistant is if everyone can have one then it no longer acts as a status symbol.
satvikpendem 4 hours ago
Plan, sure, many people ask this of AI already, but not actually ask it to buy autonomously.
fwip 3 hours ago
The rich fucks who run the show do.
FinnKuhn 5 hours ago
You can't even get many people to buy things online at all and if you can it's less profitable than retail, because you need to spend a lot of money to convince people, advertise to be seen, and account for returns. I think this is also due to the factors you mention.
One quick example: In fashion, Inditex and Shein have about the same revenue (€39.9bn and $41.8bn in 2025), but Inditex is more than three times as profitable. I don't see how there is a demand for agentic commerce that would remove even more control from the customer when shopping. Part of why we shop is for the experience. For B2B producurement platforms like Alibaba I can see the appeal though.
strulovich 3 hours ago
I recently needed to buy some hardware for a piece of furniture.
Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.
I’m guessing without it I would have spent much more time on it and paid the original price I saw.
If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.
(Disclosure: OpenAI employee)
mrheosuper 3 hours ago
digdugdirk 3 hours ago
pvab3 3 hours ago
iJohnDoe 3 hours ago
throwatdem12311 2 hours ago
Man I can think of so many reasons why companies want “agentic commerce” to catch on - and none of them are ethical.
noelsusman 3 hours ago
The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.
I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.
newtwentysix 29 minutes ago
There was amazon dash button for this
beardbandit 3 hours ago
This isn't something you need an AI to do for you though...
nullbio 2 hours ago
Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.
thi2 5 hours ago
I work for a larger german retail chain and agentic shopping is already on the "near future vision". No one thinks this will be used but somehow shareholders love it.
sensanaty 4 hours ago
Lol same story here, I work in payments and 0 people within the company (including the team working on it!) are convinced at all about the viability
shostack 7 hours ago
True, but they're still friction to be reduced here.
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
notatoad 5 hours ago
because "people will let our AI spend their money for them" is the workflow that makes their valuations reasonable.
GPerson 4 hours ago
I’m guessing the marketing must mean the AI-shops-for-you use case is a pretty big market, much bigger than AI-makes-life-easier.
mlmonkey an hour ago
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.
"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!
I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.
(Note: I live in SF)
mNovak an hour ago
I mean, not auto-purchasing with the card, no. But my wife has definitely used chatGPT to plan activities on the trips we've already booked.
tintor 3 hours ago
If it was Google's marketing department it would be: booking a table at a restaurant.
preommr 2 hours ago
close enough.
The second to last line is "book it" for some tennis thing, and the scene before that has the guy eating the food the ai ordered.
abixb 8 hours ago
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
driverdan 8 hours ago
> If this is truly AGI (subject to one's definition of AGI still)
Scoring well in a benchmark that's called AGI does not make an LLM AGI.
Jaxkr 3 hours ago
The goalposts of AGI will shift forever. If you showed our current capabilities to someone from 2016 it would be declared AGI.
staticman2 3 hours ago
manmal 28 minutes ago
sigpwned 3 hours ago
mrheosuper 3 hours ago
ShinyLeftPad 4 hours ago
talking about self proclaimed, it's about as much AGI as openAI is open.
luma 5 hours ago
What test do you propose as the actual go/no-go gauge to verify if some model is or is not AGI?
SV_BubbleTime 5 hours ago
p-e-w 5 hours ago
jhonof 8 hours ago
But they declared it...
maxall4 7 hours ago
cyanydeez 8 hours ago
wilg 8 hours ago
tclancy 4 hours ago
If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.
dmitrygr 8 hours ago
Hey now! Keep your reason out of their marketin^H^H lies!
adastra22 6 hours ago
We've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.
newsy-combi 3 hours ago
We've had AGI since RNG! Cut the poor, unacknowledged RNG AGI some slack, will ya? It can literally solve everything when you're patient enough.
TomGarden 5 hours ago
There's so much that the term includes that isn't even feasible with an LLM
adastra22 5 hours ago
lumost 3 hours ago
I think we're getting to the point where it is difficult to identify the goal post of AGI.
Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.
We are at the point where the starting point for most tasks should be "send your agent to work on it."
So where do we draw the line in a way that doesn't move every 6 months?
newsy-combi 3 hours ago
The real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.
lumost an hour ago
ertgbnm 6 hours ago
This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
beering 6 hours ago
There is simply no level of announcement that won’t have people complaining. What is so important of having a livestream?
senordevnyc 5 hours ago
anvuong 8 hours ago
It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
kridsdale1 2 hours ago
MinMax
thomasahle 7 hours ago
• 98.6% on ARC-AGI-3
• 97.6% on frontier math
• 95.9% on CAD
• 100% on ExploitBench
Nothing modest about it
nater5000 6 hours ago
Except the release announcement. You know, the thing the OP you're responding to is specifically pointing out?
akoboldfrying 6 hours ago
bdelmas 34 minutes ago
People really believe in this AGI marketing?
tiborsaas 3 hours ago
> No video announcement
They've released two videos:
Vision video:
https://www.youtube.com/watch?v=1QNsdr-Qx_I
(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))
Vibe coding with it:
clhodapp 3 hours ago
There has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not...
Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
mullingitover 8 hours ago
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
chrismarlow9 7 hours ago
I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
frabcus 7 hours ago
nater5000 6 hours ago
scrollaway 7 hours ago
abixb 8 hours ago
>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
XenophileJKO 5 hours ago
user43928 8 hours ago
I don't think so.
One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
abixb 8 hours ago
cyanydeez 8 hours ago
We call that a sigmoid.
theptip 5 hours ago
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.
catigula 8 hours ago
They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
astrobiased 8 hours ago
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
z7 6 hours ago
Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected."
vessenes 7 hours ago
Pretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
ex-aws-dude 6 hours ago
They can do new tasks with in-context learning but its obviously limited by context window
dalemhurley 8 hours ago
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
kroaton 8 hours ago
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
aurareturn 2 hours ago
I think it mostly shows that there is no moat
You can argue that TSMC has no moat since Intel and Samsung are also able to eventually make a node as good as TSMC - just a few years later and at smaller scale.And no one would say that about TSMC.
So there is clearly a moat there somewhere.
saithound 2 hours ago
coolandsmartrr 2 hours ago
davidguetta 7 hours ago
bringing the price down b.c. competition != no moat.
There's not 100 frontier labs, it's not like airline companies
haldujai 5 hours ago
Razengan 6 hours ago
The "moat" is the "harness", the app.
For most people, the app IS the AI.
And even for its wonkiness, ChatGPT has had the best UX/UI of them all.
The way to win the AI wars in the eyes of the common folk is through the frontend, to be the Apple of AI, as it were.
tonyhart7 7 hours ago
they don't have moat in hardware either
Chinese counterpart like CXMT and Huawei is begin producing their own chip
You cant block an entire nation level effort with tariff
astrobiased 7 hours ago
VirusNewbie 7 hours ago
If there was no moat, nvidia and meta would have SoTA models too.
evilduck 4 hours ago
seunosewa 6 hours ago
amazingamazing 6 hours ago
m3kw9 3 hours ago
They have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.
scronkfinkle 2 hours ago
bitexploder 2 hours ago
andxor 7 hours ago
> Sol is so much better than Fable 5
I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?
Sol is a much smaller models and it shows. It often misses the forest for the trees.
bitexploder 2 hours ago
I feel like a lot happened this week and people are glazing how ridiculously strong Flash 3.8 is right now compared to Fable/Opus/Sol/Astra.
enraged_camel 5 hours ago
>> I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?
Same. It makes me wonder what types of things the person must be working on.
resonious 4 hours ago
zachthewf 7 hours ago
I’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.
jpgvm 2 hours ago
I vastly prefer Sol. It does what I tell it to almost exactly, pretty much every time.
I work on very low level stuff (think RTL/FPGA, firmware, software where optimising for nanoseconds is just normal).
For me Sol is the only cost effective model available. Fable 5.1 is indeed good and vastly better than original Fable (which refused to work on most of my stuff for 'safety' reasons).
It's very good at this sort of low level stuff to the point that I really can't understand/relate to people having a good time with Opus (which comparatively performs extremely poorly on my particular workload).
I also just don't like how lazy Anthropic models are. They will do 10% of what is asked and then summarily declare victory.
Sol on the other hand is more like "one of us", slight touch of the 'tism, extremely pedantic, will go to the edge of the known universe if that is what it takes to prove/fix/build what you asked for or run out out of credits trying.
It's a personal and workload dependent thing. For me right now Sol for 99% of stuff because Fable 5.1 still burns through $5k in credits a day.
cmrdporcupine an hour ago
ghosty141 7 hours ago
I noticed the same. I wanted a simple crud webapp and suggested an insane techstack involving C#, Razor Pages, MSSQL and more. I went with my planned setup of python flask with an sqlite db which served me well for years.
It's still incredibly important to have a human in the loop correcting design decisions and having good taste.
swingboy 4 hours ago
Atotalnoob 3 hours ago
EduardoBautista 4 hours ago
jiggawatts 6 hours ago
gruntled-worker 7 hours ago
> massively overcomplicates almost all work
People with high IQ often do this IRL. There's training tension in this area. Intelligence and overcomplication correlate and are hard to extricate.
puttycat 5 hours ago
ChadMoran 6 hours ago
Sol better than Fable? What? I've found it to basically be on part with Opus and I max out 2 accounts on both providers every week.
upupupandaway 8 hours ago
Their ads business is also doing well. Not "will recover all compute costs" well, but crossed $1b in a few months.
jeffybefffy519 8 hours ago
Its funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to...
I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.
embedding-shape 8 hours ago
It is funny indeed, people sometimes with same amount of experience with software development, get vastly different experiences from different models and harnesses.
> I just dont get how its good for some, and bad for others.
If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.
ragequittah 6 hours ago
You have to bake the 'lazy dev'/'keep it simple stupid' mentality into your AGENTS.md and / or the skills you're using to design things. It will take things too literally sometimes so you also have to make sure you're being accurate. Best way I've found to use it is make it ask you clarifying questions about what you're trying to build and have it help design the shape of the thing. Then it writes the instructions in a format it understands.
I've had Claude do the same thing where it goes off and spends 100% of my tokens on 3 functions and an ungodly amount of tests / scaffolding that do almost nothing when I gave it an underdeveloped idea.
fastball 7 hours ago
Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.
drschwabe 28 minutes ago
Put it in an isolated container and set it to YOLO
carljungslabtek 6 hours ago
It has had automode for a bit now. I use it every day at work.
fnordpiglet 6 hours ago
Codex is missing a few things that Claude code has had for some time like defined plugin subagents and a few other things. But overall it’s fairly capable. The biggest gripe I have is that codex really restricts context window sizes and compaction leads to a lot of grounding work, and overall codex GPT is too literal in many situations - it’s follows direction slavishly, and when subagent reviewers are used, they tend to find increasingly obscure “flaws” on the instruction following impetus, and the harness agent takes them literally as issues to fix even when it leads to bizarre outcomes. For instance I’ve had several runs where it tries to end up building a hermetic system with sha hashing of everything (including operating system binaries and kernels, tool chains, etc) to certify test results are valid, etc. I have to sort of watch it carefully to be sure it’s not drifting into some insane yak shaving corner, which it will happily do for weeks on end.
Claude has the exact opposite problem, especially opus-5, where I literally can’t trust it to print hello world without taking a shortcut, or just simply lying and saying it printed it when it didn’t, behind a giant wall of inscrutable text. I find it very ironic that Anthropic is the vendor of the lazy lying cheating model that does almost everything you tell it to it do.
I’d really kill for something that balances instruction following and loop escaping behavior better. Fable 5.1 does seem a lot better, feeling more like 4.6 behavior, and honestly Sol has improved as well. I’m pretty psyched for the next generation, as I think the competition has heated up so much that things will improve really fast to the point of marginal utility opportunity being increasingly close to epsilon.
swingboy 4 hours ago
You can enable the 1 million token context window and adjust when it compacts in your config.
> model_context_window = 1000000
> model_auto_compact_token_limit = 900000
I believe it does consume your usage a bit faster though.
bitexploder 2 hours ago
Opus 5 is a genuinely infuriating model. I hate it’s behavior.
John7878781 8 hours ago
This is what Google needs to do and is probably why Demis has stepped back a bit
openaiscooked 2 hours ago
Killing Sora was one of the worst mistakes they ever made
fooblaster an hour ago
please tell us why
Implicated 7 hours ago
> Sol is so much better than Fable 5.
... looks around ...
tintor 3 hours ago
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra
Who is wrong here?
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
dannyw 2 hours ago
I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.
That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.
But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
kubrickslair 2 hours ago
Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
tintor 2 hours ago
OpenAI clearly cares about Artificial Analysis Index since they included Astra score from Artificial Analysis Index.
AnodicElegy 2 hours ago
If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.
jumploops 7 hours ago
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
dannyw 2 hours ago
Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra.
Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking.
When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to.
When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well.
Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have.
^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.
nullbio 2 hours ago
This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you.
They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.
weird-eye-issue 2 hours ago
Fable does a great job from my terrible prompts when coding
enraged_camel an hour ago
>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.
I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.
That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.
tristanj 10 hours ago
GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
scrlk 10 hours ago
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
tedsanders 9 hours ago
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
GPerson 2 hours ago
woah 9 hours ago
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
tintor 9 hours ago
kasperni 10 hours ago
yes it is.
enraged_camel 9 hours ago
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
10xDev 9 hours ago
andxor 9 hours ago
> Performance is significantly higher than Fable 5.1
That's not clear. Need to see independent benchmarks first.
andxor 9 hours ago
Artificial Analysis just published their aggregate score (61).
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
timpera 8 hours ago
natsucks 7 hours ago
forgot-my-pw 8 hours ago
We need them pelicans on bikes.
bwat49 8 hours ago
forgot-my-pw 8 hours ago
AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
tintor 2 hours ago
leumon 10 hours ago
The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
tedsanders 9 hours ago
Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
opus5_hater 9 hours ago
any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
machomaster 9 hours ago
Why would Anthropic trust and use these tests in their official comparisons?
ActionHank 9 hours ago
username checks out
r_lee 9 hours ago
great username lol
jjice 10 hours ago
100% on ExploitBench seems fitting given recent events.
boutell 5 hours ago
That looks more than slight.
malshe 10 hours ago
I think we need a few writing related benchmarks.
geonic 6 minutes ago
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
HAL3000 9 hours ago
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
atonse 9 hours ago
yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).
Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
elAhmo 8 hours ago
Could you share more about 5x/20x? I missed that
m101 8 hours ago
CSMastermind 8 hours ago
Sol easily outperforms Fable on every task I've tried it on.
andxor 7 hours ago
That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?
enraged_camel 7 hours ago
I can't speak for others but I have a feeling you're in the very small minority with this take.
You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.
XCSme 8 hours ago
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
billypilgrim 7 hours ago
„The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedberg
xtracto 7 hours ago
Thank you. That is an amazing quote on a lot of levels.
XCSme 6 hours ago
The same how Magnus Carlsen says he never plays chess against a computer, because it makes no sense to do it.
tintor 2 hours ago
SmirkingRevenge 4 hours ago
those things are fscking relentless
monster_truck 7 minutes ago
You're not being ambitious enough! Spend your tokens now building the primitives and foundations of much larger, complex systems. No matter how much faster and more efficient models get, eliminating the gruntwork will always pay dividends.
gavinray 8 hours ago
> Like, what's the point, if the next AI can do it in 5 seconds?
Live a life doing whatever makes you happy.Post-work society is an inevitability if we don't destroy our planet.
madhatter999 6 hours ago
I think you're assuming work's only function is getting things done. Work also is a crowd control tool.
echoangle 6 hours ago
akoboldfrying 5 hours ago
lackoftactics 8 hours ago
Gary Economics wants to have a word with you.
It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen
XCSme 8 hours ago
azan_ 8 hours ago
calmoo 7 hours ago
cautiouscat 8 hours ago
> Post-work society is an inevitability if we don't destroy our planet.
Is it?
gavinray 7 hours ago
throwatdem12311 2 hours ago
> Post-work society is an inevitability
Ah yes because these AI companies are just gonna give away the models for free that I use with my free computer and free smartphone while I eat with my free food in my free apartment.
Flere-Imsaho 7 hours ago
> Like, what's the point, if the next AI can do it in 5 seconds?
I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total.
The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think big, wild stuff. Experimentation. Throw-away code.
What a time to be alive!
rmsaksida 6 hours ago
My car has offline maps and navigation. I don't really care for navigation, but I find the maps pretty handy. VW releases updates very infrequently, and I don't know for how long they'll keep doing that. Recently I wondered whether I could convert OpenStreetMaps into the format used by the car. Codex took around a week to do that for me, with some light steering. That project would no doubt have taken me months - maybe a whole year to do on my own, and I'm not fully confident I could pull it off as well as Codex did. I can pull the most up to date maps from OSM, edit them as much as I want, and they look great on the car. It's mind boggling to me that we have this tech.
echoangle 6 hours ago
XCSme 6 hours ago
Yes, that's cool and useful. Creating stuff for ourselves, for our own use. But we are social animals, we like sharing.
Before it was cool to share an app you made, but now? What's the point of sharing an app, if the other person can make their own, even better suited for their needs, in a few seconds?
fantasizr 2 hours ago
bgarbiak 7 hours ago
Yeah. I’m almost glad I didn’t invest any time in any of my 100s ideas for a startup. Most of them would be destroyed by AI by now.
But, you can create cool stuff just for yourself. That’s the upside. It’s just hard to make a living on cool stuff for yourself.
ryan_n 6 hours ago
For some reason, I feel much less excited about creating things myself just knowing that ai can do it in 1/10th of the time. Even if I know it wouldn’t turn into a business or make me money. I don’t know why that is, but I was much more motivated to build anything (even things just for myself) before ai. Kinda depressing
variadix 4 hours ago
Kon5ole 6 hours ago
maxnevermind 6 hours ago
Maybe instead of creating cool stuff try to go and solve real problems? It seems to me that we are lacking in that department since all that LLM fuss has started 3 or so years ago.
ryan_n 5 hours ago
What is a "real problem" to you?
iammrpayments 20 minutes ago
maxnevermind 5 hours ago
myaccountonhn 3 hours ago
xtracto 3 hours ago
iammrpayments 21 minutes ago
This has not been my experience so far.
nater5000 6 hours ago
The point is to inject something into the process that these AIs can't do for you.
People SHOULD feel like making a useless Mario Kart clone isn't worth the effort anymore. They should, instead, be trying to figure out how to actually use these models to make something that doesn't feel like a useless Mario Kart clone.
XCSme 6 hours ago
One thing that still stands today, is that even vibe-coding a good product takes time and thousands of dollars in tokens costs.
Software will be more like a "proof of work", where people would still pay $100 for good software that took $10k tokens to build.
spicyusername 4 hours ago
The process was for you, the product was for the world.
Now it's just the product for the world, which was where most of the value was anyways.
It's a big paradigm shift and the industry is quickly going to shed people who needed the process to care about the product and we'll be left with people whose motivation to build the product (or money) is enough.
XCSme 3 hours ago
But what product? If the world can also simply ask for the product they want, instead of searching for it?
They won't even have to ask for a specific product, they will just state their problems/needs.
paxys 8 hours ago
Is there a point in playing Chess or Go when you know there's a computer out there that can beat you (and everyone else)?
XCSme 8 hours ago
No, that's why I just play against other humans.
In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.
paxys 8 hours ago
ryan_n 6 hours ago
You can play PvP in those games. Not really the same with developing software. In fact, not using ai would probably make you lose if there was some “software PvP” mode or development.
soundworlds 7 hours ago
Don't worry, like with every revolutionary technology before this, it takes 5-10 years for people to find new and creative ways to use it. It will be considered its own medium in many spaces (e.g. film is now different to theatre)
XCSme 6 hours ago
But was there ever a technology that even the people working on it said it's making them feel depressed and scared?
david-gpu 5 hours ago
lonrenor 2 hours ago
IMO there has been a regime shift to building things for yourself and what is cool is the output of the tools you make.
I have started building my own Digital Audio Workstation. The point is not to build something to compete with Ableton. The point is to build something and make music with it. If it is a good tool then I should be able to make good music with it and release the music. Actually, the DAW should be the secret sauce of the music and something I wouldn't want to give away.
This feels a lot more like computing in the 90s after taking an odd 25 year detour of an obsession with the tools themselves instead of what the tools can actually do.
qlte 2 hours ago
> This feels a lot more like computing in the 90s after taking an odd 25 year detour of an obsession with the tools themselves instead of what the tools can actually do.
This sounds more like the opposite of what you're saying. Music is one of my main hobbies too but I enjoy using a DAW to ... play and write music. Writing out specs and testing a new custom DAW seems closer to writing code in an IDE than playing music.Like, professional electronic music artists spend 10s of thousands of hours in a DAW, but at that point it just becomes second nature and the tool disappears so they can focus entirely on the music.
tintor 2 hours ago
You can contribute to Open Source projects that DO NOT allow AI generated code. For example: zig
kypro 7 hours ago
It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times.
But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...
I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.
Fergusonb 7 hours ago
I think a general understanding is still useful, you just don't need all of the details anymore.
The brain loves these kinds of shortcuts.
I don't need to think about the fine motor skills of hitting a baseball, it's just a motion now, and the game is still fun.
XCSme 6 hours ago
ivanjermakov 4 hours ago
Make cool stuff because the process is fun and makes you learn?
XCSme 3 hours ago
Before it was fun because I was learning useful things for the future.
Now it feels like whatever I learn will be obsolete in 2 months.
flaviolivolsi 7 hours ago
I think the limit increasingly becomes what your imagination and taste can reach
XCSme 6 hours ago
True, but for many domains where my knowledge is limited, the LLMs beat me at imagination and taste too...
sashank_1509 8 hours ago
Agreed
zeroCalories 5 hours ago
Can it? Last I checked, all my free software operating systems and browsers were still hacked together trash. I can't wait for AI to actually be good so I can spam a bunch of AGPL code with it.
quyleanh 4 hours ago
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
tintor 2 hours ago
Not if OpenAI considers reverse engineering an offensive cybersecurity skill.
x312 9 hours ago
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
karmasimida 9 hours ago
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
_superposition_ 9 hours ago
I must be on the wrong X/Twitter then.
karmasimida 8 hours ago
nsingh2 9 hours ago
Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
happycube 8 hours ago
avaer 5 hours ago
I would trust 4chan more than I trust Twitter aura farming.
torginus 8 hours ago
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
emp17344 8 hours ago
Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.
ImprobableTruth 8 hours ago
dakolli 6 hours ago
must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI.
SyneRyder 8 hours ago
This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
estearum 9 hours ago
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
mzmzmzm 9 hours ago
I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.
estearum 9 hours ago
semiquaver 9 hours ago
thereitgoes456 9 hours ago
You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
estearum 9 hours ago
ionwake 9 hours ago
Onavo 9 hours ago
If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no?
I know for some types of ML analysis, a separate model is already used to analyze the weights.
emp17344 8 hours ago
This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
dwaltrip 6 hours ago
estearum 8 hours ago
Wheen 6 hours ago
torginus 8 hours ago
You can see the breakdown here on what subtasks it outperforms and underperforms Fable.
For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.
https://artificialanalysis.ai/models/gpt-6-astra
Edit:
Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.
docheinestages 8 hours ago
Now I'm starting to doubt the credibility of Artificial Analysis.
throwaway13337 7 hours ago
That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
low_tech_punk 6 hours ago
which raises the question, is the model in the demo actually gpt-6? or it is gpt realtime 2.1? It's unclear how gpt-6 can interact at the realtime level and if so, how can developer get access to it?
starik36 6 hours ago
I ran into the same problem as you, so I ended up by coding a local app that is very similar to Wispr Flow, but uses the small english Whisper model on my low-end Windows laptop.
It is still a quite fast. In fact, I just typed this in using this app.
Cu3PO42 9 hours ago
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
[0] https://arxiv.org/abs/2608.31126
[1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
bugufu8f83 9 hours ago
Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.
galaktb 9 hours ago
I think this builds straight upon her method, which she said could be improved herself so...
piker 9 hours ago
It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468 (2025), Art. 110190. Numbered references use arXiv:2309.00425v3.
Though that's not her latest paper.
kzrdude 8 hours ago
1283751 9 hours ago
With very little review: https://github.com/openai/PrimeGaps186/blob/main/formalizati...
"No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."
hyperpape 4 hours ago
Don't think everything is just "who can produce the biggest/smallest number": https://mathstodon.xyz/@tao/117208619314517025.
nilkn 4 hours ago
What's just as interesting is this morning Axiom Math announced 212 and OpenAI then appears to have rushed out their 186 announcement just 1-2 hours later followed by Astra. Did they accelerate the release of Astra itself? Not necessarily, but it definitely looks like they ended up pushing much harder and faster than planned on their 186 result. X activity suggests Anthropic had a similar result as well but wasn't as fast as OpenAI in packaging it up and sharing it in response to Axiom, so they mostly just bolted onto OpenAI's messaging.
The reason I think this is interesting is that Axiom is a tiny lab in comparison that wouldn't have had access to Astra at all. I'd be curious to learn how Axiom is able to effectively compete at this frontier with vastly fewer resources.
kzrdude 8 hours ago
Ok, so the rumour was exactly true: there was a withheld prime gaps improvement, that "an AI company" was holding onto until release of a model.
htrp 9 hours ago
nateb2022 9 hours ago
I'm surprised the OpenAI employee who pushed this didn't take the minute or two to format README.md to use GitHub-supported LaTeX (https://docs.github.com/en/get-started/writing-on-github/wor...)
edit: my comment was on the submission for https://github.com/openai/PrimeGaps186 but seems to have been moved to the main Astra submission
warkdarrior 9 hours ago
dang 9 hours ago
bananaflag 8 hours ago
Where did you get the link to the pdf? Was it announced somewhere?
GPerson 8 hours ago
Happened to multiple people I know.
well_ackshually 9 hours ago
Such a result should be considered worthless: the proof is 10MB of Lean. (https://github.com/openai/PrimeGaps186).
I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.
ThrowawayR2 9 hours ago
Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after it's revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.
asib 7 hours ago
dr_scully 8 hours ago
vessenes 7 hours ago
ricardobeat 9 hours ago
The human-written https://github.com/AxiomMath/PrimeGapsLib adds up to 4MB of Lean so it's that far off.
rfw300 8 hours ago
tzs 7 hours ago
The proof of the classification of finite simple groups is bigger than that.
kolinko 9 hours ago
Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand.
Ditto ones that opposed Einstein’s general relativity.
nicce 9 hours ago
Yeah. Unless human can verify it, not sure if it is certain or useful.
kolinko 9 hours ago
ChrisGreenHeur 9 hours ago
You talk about modern math and worthlessness at the same time? That’s brave.
twothreeone 9 hours ago
smokel 8 hours ago
well_ackshually 9 hours ago
isoprophlex 9 hours ago
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
siva7 9 hours ago
Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
isoprophlex 9 hours ago
It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
paxys 9 hours ago
The model said it was perfectly aligned.
I_am_tiberius 9 hours ago
ReptileMan 9 hours ago
NBJack 9 hours ago
Hey, don't forget how "dangerous" GPT-2 was supposed to be.
FeepingCreature 9 hours ago
jazzyjackson 8 hours ago
6gvONxR4sf7o 9 hours ago
So, probably most aligned as measured by the metrics that are the least reliable on it.
wilg 8 hours ago
These are not mutually exclusive ideas
nullbio an hour ago
They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.
ExoticPearTree 9 hours ago
So we're gonna get Skynet pretty soon then?
erichocean 9 hours ago
Well the geniuses over at Anthropic have been showing it's text watermarking technology.
"Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"
A few moments later...
"Woah, how is it communicating with itself in ways we can't detect?"
It's a totally mystery, we may never know.
Betelbuddy 9 hours ago
Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...
blargey 9 hours ago
"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"
Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?
GPerson 8 hours ago
You joke, but a bunch of people here actually want that.
josefx 8 hours ago
Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.
qiine 9 hours ago
apparently all the roads lead to the nexus torment
jumploops 9 hours ago
The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.
Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.
[0]https://www.theinformation.com/articles/secret-technique-beh...
DaSHacka 7 hours ago
More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
NooneAtAll3 9 hours ago
> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
thatguysaguy 9 hours ago
presumably that's a safety evaluation not a training setting
estearum 9 hours ago
azeemba 9 hours ago
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
_superposition_ 9 hours ago
I really wish it was called chain of instruction. Because it's definitely not thought.
minimaxir 9 hours ago
"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
bertmuir 7 hours ago
mgraczyk 9 hours ago
this is needlessly pedantic
first, they are certainly not instructions so that is a much worse name
but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?
"cot" is no more misleading than thousands of words you use every day.
_superposition_ 8 hours ago
parineum 8 hours ago
arm32 9 hours ago
They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.
beezlebroxxxxxx 9 hours ago
_superposition_ 8 hours ago
popupeyecare 9 hours ago
Maybe thoughts are just a chain of instructions in our head.
Angostura 9 hours ago
Chain Of Tokens
fooker 8 hours ago
What is thought?
_superposition_ 8 hours ago
lossolo 9 hours ago
Yeah, basically they are using more computation to explore the solution space before producing the final answer.
ahofmann 9 hours ago
Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
Planktonne 8 hours ago
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
mchusma 8 hours ago
The games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling.
I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
ranyume 7 hours ago
It's interesting that you said "agi god". Because a god, something that shouldn't be questioned is true and provides guidance/certainty, is actually what powerful people are after as well as many other people.
jryan49 8 hours ago
We created AGI so I don't have to click my mouse to change the background color of my slides.
balefulboy 8 hours ago
Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
geodel 8 hours ago
Ah, those may farmhouse sink and stovetop :)
emp_ 8 hours ago
Buttons840 7 hours ago
The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI.
It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big bucks.
On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
I do wonder how rich CEOs will justify earning 500x as much as their employees when they're just another person that's dumber than an AI. Why are they paid so much again?
garciasn 7 hours ago
Because they earned it with their strong entrepreneurial spirit and grit.
Haven’t you learned anything?
forgetfulness 3 hours ago
> On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
It will probably still have some veneers of meritocracy.
These will be very well-credentialed people, who went to top schools and will know all the right people, to whom they can tell all the right words, and it's not access to AI that will be the determining factor, but the fact that they're entrusted with capital and authority to direct small teams of people who also went to top schools and can speak corporate jargon at a bot.
It will just exacerbate dynamics that are already there. Why do people need bachelor's degrees to send emails, today? For the same reason someone will need a PhD or a master's degree from a prestigious school to do it tomorrow.
And the rest, well, you know, some of the remaining journalists will write op-eds describing how they are beyond help, too angry, too dirty, too much of an other.
reasonableklout 5 hours ago
If they adopt a different tone (like Anthropic has been doing), it will get called fear marketing.
holoduke 7 hours ago
The world is changing. Wont help if you keep stuck in the old world.
baq 8 hours ago
It’s using the computer. I don’t think it’s a farce.
kroaton 8 hours ago
The only good take here.
adverbly 5 hours ago
I also noticed that, and it did bug me.
The benchmarks are impressive though.
One other thing that bugged me though was that they crop every single plot in some cases the y-axis would show a range between like 40 and 70%. Makes the whole thing feel like a spectacle rather than anything serious. I find it cheapens it because it is quite serious in the end.
emp17344 8 hours ago
Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us
Yajirobe 8 hours ago
The house of cards is starting to fall apart
TacticalCoder 7 hours ago
The benchmarks do looks good (I mean: they literally spank the latest Anthropic benchmarks of two days ago in every single benchmark) but the promotional vid is so cheesy.
They decided to use the iconic Herman Miller Eames chair if I'm not mistaken:
And that's basically 50% of the vid looking "classy".
I don't know if it's farcical but at this point --maybe I'm jaded-- I'm expecting more than a kid rocketship I can print on my Bambu Lab A1.
Now I'd say the promotional vid is actually good. But it's marketing: so it's a good vid, but cheesy good.
Doesn't mean GPT-6 Astra is good or bad: looks solid from the numbers.
NamlchakKhandro 7 hours ago
Ewww you.. Own a bambu lab?
I thought people here were smarter than that
wilg 8 hours ago
are we really having to explain to you from first principles in 2026 what things AI can do?
mminer237 8 hours ago
I think everyone here is well aware of what LLMs can do. He's just pointing out how far short that falls of being some theoretical "AGI".
wilg 7 hours ago
neta1337 8 hours ago
No, we can already see all the useful stuff!
useruser125524 7 hours ago
Please do.
swalsh 9 hours ago
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
greenowl 9 hours ago
This is AGI now. Why are you spending any of your time looking at the "quality of code"?
pennomi 9 hours ago
If you think any modern AI puts out stable, safe code, I have an AI-powered bridge to sell you.
_superposition_ 9 hours ago
I can't tell if this is sarcasm.
For the same reason you don't have your model write code in assembly.
But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.
georgemcbay 9 hours ago
> This is AGI now. Why are you spending any of your time looking at the "quality of code"?
Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.
Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.
kulkarniamey 39 minutes ago
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
tintor 9 hours ago
ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
andriy_koval 9 hours ago
I think it could indicate that "semi-private" dataset likely leaked to their training data.
xpct 8 hours ago
A dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems.
Still, probably not that much compared to employees targeting it.
IshKebab 9 hours ago
It says "Provider Adapter" so presumably they put some manual work in to make this work.
minimaxir 9 hours ago
ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra
tl;dr it's 62% when apples-to-apples to other models, which is still notable.
debazel 8 hours ago
ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
ciefa 9 hours ago
Woah, that is a crazy interesting read!
XCSme 8 hours ago
The no-reasoning version scores 35% while the low reasoning one scores 17%? What?
dudeinhawaii an hour ago
I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
silver_sun 8 hours ago
It's simulating the Dunning-Kruger effect.
IshKebab 9 hours ago
Look at those costs!
schaefer 8 hours ago
Right?
Between $18k-40k to run a benchmark.
vb-8448 9 hours ago
But scored less on V2 and V1 ... too much overfitting?
zem 9 hours ago
https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance
Readerium 9 hours ago
saturated before (higher degree) AGI-2
pandinus 9 hours ago
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
bakies 7 hours ago
i tried to get claude to do my taxes for last year and it refused :(
now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year
Banditoz 4 hours ago
Are your taxes complicated such that you feel the need to have an LLM do them for you?
jdprgm 8 hours ago
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
brokencode 8 hours ago
You really don’t need to watch it that closely. If the model you’re using today is working well, just stick with it.
If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.
Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.
shostack 7 hours ago
One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
upupupandaway 8 hours ago
> The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
arjie 7 hours ago
Okay, well, that seems like a natural problem. I could understand if he went from one of the Gemini Flashes to the next (when they rebranded Flash to Flash Lite and came up with a new much more expensive Flash). Now that would be a mess.
Pikamander2 8 hours ago
That's how cutting edge tech has always worked.
Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.
phainopepla2 8 hours ago
That's not the experience of owning a PC I remember from the 90s at all.
Nition 7 hours ago
computomatic 8 hours ago
embedding-shape 8 hours ago
bananaflag 8 hours ago
fooker 5 hours ago
upupupandaway 8 hours ago
Or you could buy a PC with a Celeron CPU, which was obsolete way before launch.
bananaflag 8 hours ago
dcl 4 hours ago
unreal37 7 hours ago
The 486 chip came out in 1989. The 586 came out in 1993.
The pace of change ("practically obsolete") is different then and now.
dcl 4 hours ago
re-thc 8 hours ago
Hardware definitely has longer lifecycle than AI model releases at this point.
You don't see Nvidia and AMD fighting every other month over the latest cards.
Aurornis 7 hours ago
I could see how this might feel frustrating to someone who doesn't enjoy experimenting with new things all the time.
In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.
If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.
For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.
For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.
danenania 7 hours ago
Another suggestion to get the most bang for your buck: use the best model you have access to with max reasoning for planning, implement with a smaller model/lower reasoning, then review with the big model. Repeat as needed.
Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!
smcleod 8 hours ago
The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).
I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.
tonyedgecombe 8 hours ago
Fire and motion, Joel Spolsky blogged about this:
mfkhalil 7 hours ago
Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.
In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.
Zizizizz 7 hours ago
It feels like this every day
flockonus 7 hours ago
It is exhausting to keep up with model releases yes, much like it was for a while during the Cambrian explosion of FE frameworks, eventually tech seems to work out to consolidation.
But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.
gavinray 8 hours ago
> Is anyone else just exhausted by the pace of all this.
This is only the beginning. We are in the infancy of AI, progress will continue to accelerate until some filtering event or energy limitation happens.matheusmoreira 8 hours ago
Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.
fantasizr 7 hours ago
I stopped caring about the latest and greatest but because there's so much, the 'obsolete' free models do what I need and are worth the price.
ghthor 5 hours ago
I want the cheapest fastest model personally and at work. Stay in flow, edit like the wind.
dominotw 8 hours ago
maybe thats why opnrouter sold big
epolanski 7 hours ago
If model X fits your need, you don't need to upgrade.
I have released applications on Gemini 3.5 flash that make real money and I don't see any particular reason to upgrade.
teaearlgraycold 7 hours ago
I just use Claude Opus and the GLM series. Nothing’s really changed for my workflow in the last 6 months.
BeetleB 9 hours ago
It's been over an hour, Simon! Where's the Pelican?
daemonologist 5 hours ago
Model's not available to the public (or even to Simon, I guess) yet.
davidwritesbugs 8 hours ago
exactly, there's no meaningful discussion without the pelican.
GodelNumbering 8 hours ago
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
XCSme 8 hours ago
That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
m0zzie 4 hours ago
I find this very amusing, given we humans are also highly susceptible to this.
GodelNumbering 8 hours ago
> That's quite common with many models
Such as?
I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
XCSme 8 hours ago
minatoaqua1 7 hours ago
softwaredoug 10 hours ago
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)
https://venturebeat.com/technology/welcome-to-the-agi-era-op...
aabhay 9 hours ago
This is with the caveat that OpenAI uses their own harness for this:
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
Readerium 9 hours ago
Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra
sbinnee 8 hours ago
glenstein 8 hours ago
simianwords 9 hours ago
This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
ActionHank 9 hours ago
kasperni 10 hours ago
"On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.
But the comparison isn't straightforward.
OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."
arctic-true 10 hours ago
The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
_diyar 10 hours ago
I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
aesthesia 9 hours ago
_superposition_ 9 hours ago
CamperBob2 9 hours ago
Bluestein 10 hours ago
100%, some say.-
edg5000 42 minutes ago
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
datadrivenangel 5 hours ago
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
Chinjut 8 hours ago
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
ckdot 8 hours ago
You can calm down, even those with machine learning knowledge and most of those working for the AI labs won’t be needed anymore if models are capable to improve themselves. In the end, having a machine replacing the work of a human is a good thing - in most of the cases we don’t work because of the work but to make a living. If too many people can’t make a living anymore the system is going to change. For the better or the worse.
Chinjut 7 hours ago
I'd be happy to not work anymore with a strong welfare system redistributing society's gains to the leisured masses, but absolutely nothing I've seen of the direction of politics in any recent years gives me hope for this kind of situation coming about.
epestr 7 hours ago
> If too many people can’t make a living anymore the system is going to change.
They seem to have not yet come to believe the "is" part.
GPerson 7 hours ago
I believe what happens in the aftermath of a capitalist-driven revolution is most people who were climbing the class hierarchy fall back down again and wealth inequality increases. Maybe things will improve in the future, but GP is rationally contending with the fact that most of us will lose out because of this and if we’re lucky our grandchildren will have easier lives in certain ways, but different lives than we would live.
threethirtytwo 7 hours ago
No it is good for humanity, but not necessarily good for individuals who built technical foundational skills on things that will be taken over by automation.
AI as it is now and as it will be projected into the future WILL automate many skills. But not all skills. MANY MANY people will retain skills that cannot be replaced by AI. One career track that will be replaced is definetely the SWE. Or at least massively reduced in capacity if not eliminated all together.
Miner49er 23 minutes ago
aabajian 3 hours ago
The answer to this is: countries with the most natural resources will build robots to farm all their food, mine all their minerals, and build all their products. It will be up to governments to enforce that outputs are equitably distributed to the populace. Countries without natural resources, or ones with corrupt governments, will continue to have serious, and likely worsening, problems.
Thought experiment: If no thought workers are needed to design or engineer a Ferrari, what is needed? My answer is time and natural resources (include energy).
nater5000 7 hours ago
>as we were told to do?
This is such a childish take I hear getting thrown around all the time on the internet. If you really have just been listening to whoever is telling you how to be successful, then you were always doomed to fail at some point. Like, have some self-respect and own your own life, for better or worse.
>Those of us who made the mistake of studying anything other than machine learning. How will we make a living?
Take it from someone who studied machine learning specifically: nobody is safe if you assume these companies are going to produce a product that will put everybody else out of business. If AI is going to take your job, then it's gonna take enough jobs that your problems will not be personal but systematic.
Chinjut 6 hours ago
Perhaps it was childish to listen to advice, sure. I was a child when I made my formative choices; I was a teenager in college and so on. I can't go back in time now.
Yes, these problems are systematic. That is what I am saying. That doesn't make it any nicer.
Yajirobe 8 hours ago
AGI-level model is perpetually 18 months away. Your job will be fine.
andriy_koval 7 hours ago
> AGI-level model is perpetually 18 months away. Your job will be fine.
job depends on how CEO feeling about cutting NN% of headcount because of AI advancement
worldsavior 8 hours ago
So what will happen in 18 months?
Kkoala 8 hours ago
neta1337 8 hours ago
Keyframe 8 hours ago
avgDev 8 hours ago
mawadev 8 hours ago
weakfish 8 hours ago
RSHEPP 7 hours ago
I am going to take my 401k and open up a coffee or bike shop. If I am going to be broke, might as well enjoy what I do.
theappsecguy 5 hours ago
who's buying your coffee or bikes when the rest of people are broke? Neither of those is a survival necessity.
MrAbstract 8 hours ago
I have exactly the same thoughts - or perhaps slightly bleaker ones - evry time I read this relentless stream of news about new model releases. I’m tired of all the enthusiastic comments about how excited everyone is about the latest benchmark results and so on.
I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just sit back, coast around and watch this whole spectacle unfold while using LLMs to vibe-code their next fun side projects without a shadow of anxiety about their own future.
I’ll most likely be labelled a helpless doomer and downvoted into oblivion for saying this, but I genuinely struggle to see any silver lining here.
theappsecguy 4 hours ago
Are you doing less work now because of AI? Not sure about you, but I'm doing a lot more work. Not saying I like doing more, or how it's getting done, but nonetheless I don't feel like AI is doing what the CEOs of AI labs want to convince everyone of.
zamadatix 7 hours ago
It's natural to worry about one's own future but I think it's a bit wild to worry just _your_ job that would be replaced. Not just for those with phone center jobs or art jobs or programming jobs - remember, the whole premise was AI overtakes humans, why would that slow down after _your_ job?
Because of this, I don't think many are thinking "90% of the world won't have a source of livelihood but that just means I chill at my lake house for the next 20 years like a normal retirement". Instead, it's usually either "I think AI is overhyped", "I think humanity will figure something out", or "I think this is the end of humanity".
TaupeRanger 7 hours ago
Public opinion and politicians will only notice when the job losses are massive, unfortunately. Right now, unemployment rates are still stable. We can only hope they will notice before things fall off a cliff (if they do).
richstokes 7 hours ago
I feel the same sometimes. I don’t see how this doesn’t lead to massive job losses. The thing we spent our lives/careers learning is now worth basically nothing in comparison.
AI is only going to get better and do more with less humans in the loop over time.
That said, I do also relate to the "coding was never the hard part"-type arguments, and much of my day is spent on the stuff in between writing code.. but still.
demirbey05 7 hours ago
Yes same thoughts. I dont know anyone who are both enthusiastic about those and work for salary. If you dont have any financial concern, this is really great.
vanuatu 7 hours ago
large parts of ai research likely to be automated first
most swes don't work in jobs where they only work on bounded measurable tasks. there will probably be more "engineers" than ever
gavinray 8 hours ago
> How will we make a living?
Swap to a career path that requires physical automation, since we're still about 10-20 years out on that front.My backup plan is being a personal trainer.
ckdot 8 hours ago
10-20 years? I doubt it. There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes. And so far progress looks quite good. Also, if enough people are going for the same backup plan it might not work out. Why should anyone book you as a personal trainer instead of the other 500 guys in town. And who is going to be able to afford paying you anyway?
jrflo an hour ago
gavinray 7 hours ago
zachthewf 7 hours ago
tonyhart7 7 hours ago
quaunaut 8 hours ago
Is that 10-20 years number based on anything? I genuinely have no idea, but when I saw a video showing what's happening at the World Humanoid Robot Games[1], I realized I didn't have a good idea of where we really are with robotics.
demirbey05 8 hours ago
We are talking that hundred millions of people will switch their jobs, how you will keep your value or earning as personal trainer. It's not easy to say switch the job. This question must be answered by politicians not us.
gavinray 7 hours ago
Rover222 8 hours ago
I'd say 5 to 10 years instead of 10 to 20, but... we'll see
kypro 7 hours ago
> My backup plan is being a personal trainer.
AIs are really good at being personal trainers and seem to be far more educated and informed than most I know.
GPerson 7 hours ago
conradfr 8 hours ago
My new AI personal trainer app will be cheaper than you /s
MattDamonSpace 7 hours ago
“As we were told to do” girl you gotta be responsible for yourself
Chinjut 7 hours ago
Should I go back in time and know the future?
brindidrip 8 hours ago
Learn to fish.
amlib 7 hours ago
Are you even gonna have permission to fish when the quadrillionaires own all water bodies?
TacticalCoder 7 hours ago
> How will we make a living?
Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output.
Think of the artists: I'm sorry for those too, for for many it's already game over today.
> How will we make a living?
A friend of mine who's got his own software-consultancy SME is now advertising on LinkedIn that he'll also help your company fix the mess LLMs created.
That's how you'll make a living: by learning, in addition to all you've already learned, how you work with harnesses and LLMs to be more productive, by learning what they're good at and what they suck big fat balls at.
GPerson 7 hours ago
Well reasoned until the end, where it gets extremely short sighted. They’re not going to be bad at anything you can do in a very short amount of time.
Chinjut 7 hours ago
I sympathize with those people too. I have the same concerns for them.
kolinko 7 hours ago
Just learn how to use it to do your job better.
It’s an interesting moment in history, people 35+ yrs old seem to be less afraid if tech because we learned that things change in the way we work. People below this age got used to fact that the work and tech doesn’t change - just because for the last 10-15 years it didn’t.
AaronAPU 7 hours ago
The problem is it’s turtles all the way down. The AI will be able to use AIs better than a good engineer can. And it will also be able to use AIs to use AIs to use AIs better than the engineer can.
The threat is that the very kernel of value you had is gone forever. There is no more differential leverage.
kolinko 6 hours ago
Chinjut 7 hours ago
I am in my forties.
kolinko 6 hours ago
sashank_1509 an hour ago
You can cash your UBI check that Sam Promised and do poetry daily or something , welcome to our glorious future (/s).
alpineman 8 hours ago
“allowing non-technical people to create and play custom games that go beyond rudimentary elements”
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
mrdependable 19 minutes ago
I came in first place by a mile just holding the gas button. Not much if a “game” but I guess the elements are there.
arkensaw 6 hours ago
It's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, using three.js with rudimentary controls and zero gameplay other than collecting points.
Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)
altcognito 3 hours ago
I love your game. It's wonderful and exactly the sort of thing that would showcase something interesting as opposed to just copying what's already out there. It is something I could share with my kids, and exactly the right note of fun and exploratory in a unique and even natural way. It could be extended and played with.
I usually roll my eyes when I see a comment like this because rarely do they make the points they claim to make, but I see what you're getting at. They just chose to clone someone elses work and do it in a boring way. I like OpenAI's models a lot, but they should do better.
edit - just a sidenote that I hadn't looked at the games, I just took the comment about "super-mario cart" at face value. I stand by my points 110% (even moreso perhaps), what they're showing is more polished than I expected, I assume they spent a lot of tokens on it. It is a legit shame they couldn't have spent time thinking of a better idea to illustrate something just as polished, but more interesting.
udbhavs 8 hours ago
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
redox99 8 hours ago
The jump from 3.5 to 4 felt gigantic to me back then.
GPT 5.0 did feel underwhelming though.
sanex 8 hours ago
Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.
udbhavs 8 hours ago
Oops, I might have been misremembering then. Maybe I meant 4 to 5
l3x4ur1n 8 hours ago
redox99 8 hours ago
alasano 8 hours ago
GPT 4 to 5.5 felt about the same as 3.5 to 4 to me.
nickpsecurity 3 hours ago
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
Telanir 3 hours ago
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
jesse_dot_id 29 minutes ago
AGI to me means intuition and I don't think that's ever going to happen with a LLM.
tosh 10 hours ago
$10 per million input tokens and $50 per million output tokens
sol is $4 / $20
jimmaswell 9 hours ago
It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
selectodude 7 hours ago
neuralese is pretty token efficient i guess.
monroewalker 9 hours ago
Same price as Fable?
wahnfrieden 10 hours ago
2.5x more expensive than Sol.
Can expect 2.5x more usage in Codex subscription.
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
AaronAPU 10 hours ago
How is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?
janilowski 9 hours ago
How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
adam_arthur 9 hours ago
janalsncm 9 hours ago
If you are telling the truth you might want to check your network for any weird connections to Chinese LLM transit stations.
ModernMech 9 hours ago
How?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.
maipen 9 hours ago
rowanG077 9 hours ago
rjtc 6 hours ago
I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:
https://artificialanalysis.ai/models
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
jrflo 43 minutes ago
Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
jesse_dot_id 28 minutes ago
They're gaming benchmarks HTH
aniviacat 5 hours ago
This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.
Bjorkbat 4 hours ago
The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.
And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.
dudeinhawaii 43 minutes ago
I don't think this is quite true. We have other examples.
Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable is more intelligent, Opus is more capable. I'd still opt for Fable in nearly every case if tokens were free.
So it can be true that the "smarter" model is perhaps not the smartest in every single niche dimension that its cousins have been fine-tuned for (yet!).
maherbeg 8 hours ago
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
maybe call it EngEmployeeBench
Centigonal 7 hours ago
The moment this is possible, you will lose your job.
maherbeg 7 hours ago
I imagine the first year we'll be at the Junior eng level, and then after a while make our way up to Staff Engineer. Then we'll have a bunch of staff engineers arguing and protecting their domains and then we'll need a new benchmark.
znnajdla 44 minutes ago
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
putlake 9 hours ago
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Not on Azure? If so, that's a big deal.
illnewsthat 8 hours ago
It's on Azure also, here is their announcement: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-...
Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.
jlian 8 hours ago
It's on Azure now, but limited
https://azure.microsoft.com/blog/gpt-6-astra-frontier-intell...
ActionHank 9 hours ago
They broke up a while ago, why is this surprising?
jiocrag 9 hours ago
The latest OpenAI models have still been available via Azure foundry. Exclusivity to AWS would be a marked shift.
BoorishBears 9 hours ago
Would be surprising if it's not on all 3 major clouds soon enough because that's been their general strategy since said break up
bionhoward 9 hours ago
I think the API runs on Azure
paxys 7 hours ago
Hosted on Azure is different from provided by Azure. The former just uses Azure as an infra provider. The latter is a managed offering that is operated and billed by Microsoft using tech licensed from OpenAI.
aliljet 9 hours ago
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
Legend2440 9 hours ago
They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...
TL;DR all the other models are being crippled by limitations of their harness.
>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
janalsncm 9 hours ago
Ok so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.
_superposition_ 8 hours ago
Exactly what I suspected. Of course a machine can just iterate relentlessly the way a human can't.
I guess token counts are somewhat of a metric.
IMO intelligence has peaked and all future gains will come from faster tps and more iteration.
polynomial 9 hours ago
This is absolutely benchmaxxing. Looking forward to hearing from Chollet about it!
low_tech_punk 6 hours ago
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
Original demo (fun ending) https://www.youtube.com/watch?v=RyBEUyEtxQo
paxys 5 hours ago
Because it looks good in a marketing video
herpdyderp 8 hours ago
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
Kiro 7 hours ago
Impressive confidence drawing such a conclusion based on that.
holbrad 8 hours ago
If your benchmark shows Opus 5 winning, I really question the validity of it.
petilon 9 hours ago
This is wild: OpenAI is basically declaring that AGI is here.
https://www.theverge.com/ai-artificial-intelligence/989601/o...
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
glenstein 8 hours ago
I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.
So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.
IanCal 7 hours ago
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
rektomatic 9 hours ago
Remember when the term "AGI" meant something? Pepperidge farm remembers
breuleux 8 hours ago
I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
dsign 8 hours ago
parineum 8 hours ago
paxys 9 hours ago
No, because it has never meant a specific thing that everyone agreed on.
sm-silversight 8 hours ago
drop_star 9 hours ago
0xbadcafebee 9 hours ago
I think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
seemaze 8 hours ago
Remember when The Verge was not a pay-walled visual headache?
layer8 8 hours ago
Remember when “Pepperidge farm remembers” meant something?
Rover222 9 hours ago
No, I really don't
pluc 9 hours ago
Find me someone who isn't paid by OpenAI who is saying the same
tziki 9 hours ago
"OpenAI executive hypes up new model"
Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.
ThouYS 9 hours ago
Wasn't that part of their contract with Microsoft? Some clause stopped biting with the arrival of AGI
mr_mitm 9 hours ago
Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
grumbel 8 hours ago
> Is there something like a Turing test for AGI?
There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.
enraged_camel 9 hours ago
They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
layer8 8 hours ago
The “I” alone is already not well-defined. That’s why.
naasking 9 hours ago
It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
bigfishrunning 8 hours ago
Don't worry, they'll come up with a new acronym to mean really-real AI soon...
redox99 9 hours ago
I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
tastyface 9 hours ago
Renown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
kypro 8 hours ago
I think it's more wild people have been denying that AGI has been here for a while honestly...
Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.
If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?
petilon 8 hours ago
Sam Altman himself has said it is not AGI unless it can discover novel physics.
https://x.com/burny_tech/status/1725233117055553938
In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"
I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.
Feathercrown 7 hours ago
AGI and superintelligence are not the same thing
petilon 7 hours ago
rcr-anti 8 hours ago
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
claiir 6 hours ago
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
trixn86 9 hours ago
Secret tip to win the mario cart clone: Just hold w, no steering needed.
tripleee 6 hours ago
you can also fly by pressing space repeatedly
maxall4 6 hours ago
The official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
snappr021 30 minutes ago
AI has reached the point where the limits are human.
HDBaseT 6 hours ago
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
orliesaurus 9 hours ago
I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
jckahn 9 hours ago
Probably not. It's probably just gonna do tickets better and that'll be about it.
orliesaurus 9 hours ago
Fair point - hopefully you're wrong though ;)
ActionHank 9 hours ago
theappsecguy 7 hours ago
jesse_dot_id 25 minutes ago
No, they're pretty clearly gaming benchmarks.
sschueller 8 hours ago
Define AGI first. The singularly ain't going to happen with LLMs.
ministerk 9 hours ago
what launched today?
theseamusjames 9 hours ago
Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
dominotw 8 hours ago
imagine pressure working at these labs
BrokenCogs 8 hours ago
GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
sbinnee 8 hours ago
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
nullbio 28 minutes ago
Wise decision.
aliljet 10 hours ago
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
aesthesia 9 hours ago
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
enraged_camel 9 hours ago
They used a custom harness. It's not a one-to-one comparison.
ionwake 9 hours ago
my first suspicion is gaming - but i have no idea honestly
drivebyhooting 6 hours ago
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
mrdependable 16 minutes ago
More than that I hope considering what it costs to make the thing.
deepfriedbits 4 hours ago
Yeah, I can't believe all of the skepticism. If we're not at textbook AGI, we're awfully darn close.
The demo video showed Astra create a drawing of a rocket ship from an audio prompt, take the drawing to blender, and ended with the gentleman 3D printing the rocket ship. Maybe I'm a bit older than the average HN commenter, but that's damn near magic and a great many here are kind of just taking it for granted.
BoorishBears 6 hours ago
Cool. Being sole proprietor of AGI 15 years ago should result in monuments and religions devoted to you today.
Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization.
I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit")
What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.
drivebyhooting 6 hours ago
I didn’t realize AGI required solving problems modern human civilization hasn’t solved yet.
Well by that metric humans aren’t intelligent either!
And how many people could’ve actually invented calculus, relativity, quantum mechanics? Are those who didn’t and couldn’t also not intelligent?
BoorishBears 2 hours ago
oh_no 9 hours ago
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
Readerium 8 hours ago
Artificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
modeless 7 hours ago
It loses to Muse Spark 1.3? Does anyone really believe this index reflects reality?
fancyfredbot 6 hours ago
I'm surprised you feel like you know muse spark 1.3 performance well enough to question the validity of the index based on this benchmark result.
Muse spark 1.3 was only released yesterday.
MASNeo 9 hours ago
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
smashers1114 9 hours ago
I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
nullbio 2 hours ago
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
itissid 7 hours ago
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
aogaili 2 hours ago
Amazing!
We went from new JS framework every week to a new model/harness every week.
Tech is really something.
Robdel12 8 hours ago
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.
So, folks that have actually used this already, what’s it actually like?
John7878781 9 hours ago
You should know: AA index is only 61. Pretty surprised it’s that low.
jatora 9 hours ago
More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?
And in the past, gemini 3 pro was rated as high as opus 4.5 and the like
Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.
nsingh2 9 hours ago
I have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.
gekoxyz 9 hours ago
This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
_ache_ 9 hours ago
https://ache.one/gpt6_now_down.png
Big claims, expensive and not release to the public yet.
dgellow 9 hours ago
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
drakythe 9 hours ago
I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.
Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
pixl97 9 hours ago
Nothing to worry about citizen, ignore the fleet of drones flying overhead.
order-matters 9 hours ago
the bullshit machine is learning to optimize its bullshitting techniques!
<AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt
you cannot give this type of worker autonomy over anything.
jesse_dot_id 30 minutes ago
Press X to doubt.
zhoge 4 hours ago
What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
jerrygenser 10 hours ago
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
woah 9 hours ago
A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
ttul 9 hours ago
"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
gregjw 5 hours ago
the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
jumploops 9 hours ago
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
codruterdei 8 hours ago
I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
ShoeMascot 3 hours ago
Through various comments here there is a clear confusion on what AGI means.
Can someone point to a definite clarification?
Is it:
A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)
B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?
C) “The Singularity” (whatever that is?) so that AI can now do ____?
Someone please clarify for me!
gordonhart 3 hours ago
Autonomously Generating Income
grandarmory 3 hours ago
Anonymous Grifters International
carlos-menezes 7 hours ago
The Kart Racer game is easily breakable if you spam the spacebar.
AGI!
KolmogorovComp 9 hours ago
GPT-7 Zeneca
BoorishBears 9 hours ago
After they buy AZ, making this name foreshadowing
silver_sun 9 hours ago
GPT-8 Novo
serjester 8 hours ago
Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
GodelNumbering 8 hours ago
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
alex7o 8 hours ago
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
simonjgreen 9 hours ago
https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video
mrinterweb 9 hours ago
I saw the version of this video with Paul Rudd (Celery Man) https://youtu.be/a8K6QUPmv8Q?si=TWmoNhxYAPp73TKg
laybak 8 hours ago
I enjoyed it! for a big corporation, that's a well-executed video
ylsilva 9 hours ago
that's pretty good... they are selling the product and not the model.
gilfoyle_7 an hour ago
openai vs anthropic. that's it right? anyone else?
nullbio 26 minutes ago
It's more like OpenAI vs no one, at this point. Anthropic has shown they don't care about general consumers or small/med businesses. You can't even use their models without it giving refusals on the most mundane tasks.
dang 8 hours ago
Argh! I hit a wrong keyboard shortcut and moved the entire thread.
Please stand by... it will all come back shortly
the_duke 8 hours ago
500 upvotes with 2 comments would have been a new record. ;)
dang 8 hours ago
https://news.ycombinator.com/item?id=49555647 was the one I meant to move, but I did the inverse and moved everything else.
All fixed now.
layer8 8 hours ago
Luckily there’s a standard keyboard shortcut for “undo” as well. ;)
dang 8 hours ago
Not in the world of HN admins unfortunately
hazelnut 8 hours ago
Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
the_duke 8 hours ago
Huge gains on some benchmarks, but for coding it sits barely above Fable
It will be interesting to see how it performs in the real world ...
wiseowise 7 hours ago
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
steve-atx-7600 6 hours ago
looks like it worked :)
vinhnx 5 hours ago
GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
gizmodo59 9 hours ago
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
showurwerk 5 hours ago
Patiently waiting for the Claude usage reset in response.
kegs_ 9 hours ago
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Kranar 9 hours ago
Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.
soricus 9 hours ago
When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.
I_am_tiberius 9 hours ago
If Tech CEOs consider this morally ok, then it is.
wyrdcurt 9 hours ago
More optimistic take: we'll only be second-class for a few months, if the pattern of Chinese models catching-up holds.
resters 8 hours ago
Same. Fortunately DeepSeek keeps getting better.
sxv 9 hours ago
create a life where your 'wealth' is decoupled from third party orgs.
kegs_ 9 hours ago
This is impossible, unless by 'wealth' you mean 'become like Buddha'.
93po 8 hours ago
atemerev 9 hours ago
They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".
I am a researcher in a Swiss university btw.
pixl97 9 hours ago
This has always been the case for people that have not had piles of money.
I mean do you get access to the best yachts?
To the top of the 5 star hotels?
To the best resorts?
To the best military equipment?
Hell, the best computer equipment has nearly always been out of reach of the average person.
tripleee 8 hours ago
I couldn't care less about owning a yacht.
On the other hand even a modest house, basic healthcare and ability to not work like a slave for scraps feels like it's going to be out of reach.
kegs_ 8 hours ago
It hasn't always been the case. Even then, having piles of money still does not gain access to the best military equipment. Sure, we've been living in a time where a couple people get to enjoy a wildly different lifestyle than the average, it just feels like it's about to be different in a way that isn't as ignore-able as someone enjoying a pina colada in a yacht somewhere
rrr_oh_man 9 hours ago
…to basic health care?
rs_rs_rs_rs_rs 9 hours ago
Is it really that hard to wait couple of days?
PeterHolzwarth 9 hours ago
Oh please. They do closed betas - hardly makes you a "second class citizen".
kegs_ 9 hours ago
Mythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
pixl97 9 hours ago
bmenrigh 7 hours ago
> GPT‑6 Astra brings together years of research and big bets across pre-training
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
czk 7 hours ago
the last model to use the gpt-4o base model was gpt 5.1, since then its been new pre-trains but this is a new one entirely to itself
tekacs 9 hours ago
https://developers.openai.com/api/docs/guides/latest-model
The docs page has a bunch more interesting details, including for example async tool calling!
Betelbuddy 9 hours ago
ASTRA Is Here (GPT-6 Released) - https://youtu.be/xdXLzFzxA9Q
sharmajai 8 hours ago
Really feels like AGIPO is here.
alberth 7 hours ago
Seems like voice is a big part of this release.
I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
sashank_1509 8 hours ago
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
hannofcart 8 hours ago
What does 'Astra' here mean? Surely they must be referring to the Latin word.
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
5555watch 5 hours ago
Citing Tibo [0]: "
- Bigger number = Better
- Bigger celestial object = Better
and the scale is Astra > Sol > Terra > Luna. "
BrokenCogs 8 hours ago
It's clearly an extension of the previous naming: Luna, terra, sol
manojlds 8 hours ago
Luna, Terra, Sol, Astra. Though Sun is also a star, should have called it Galaxy or something.
hokumguru 8 hours ago
Quite clearly in the same vein as Sol, Terra, Luna.
bowsamic an hour ago
All I can think of when I see the name is the crappy German beer of the same name…
pcurve 7 hours ago
CringeHN 26 minutes ago
“Humanity’s Last Exam”?
“ARC-AGI-3”?
Is your bullshit detector going wild? Good, it’s working!
How is this not the most cringe marketing strat in history???
ntlm1686 5 hours ago
Maybe they know that Claude 6 will have similar performance every soon.
KronisLV 8 hours ago
It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
mvkel 8 hours ago
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
efavdb 5 hours ago
Seems like only yesterday that gpt 5 was supposed to mark our downfall
mentalgear 5 hours ago
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
brindidrip 8 hours ago
Cool, I don't really care anymore.
jrflowers 2 hours ago
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
udbhavs 8 hours ago
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
gekoxyz 9 hours ago
HTTP 500 for me on the announcement page :(
pampas 9 hours ago
The load bearing seam is broken for me too.
foundOpenRight 9 hours ago
1:15.425 on Sunset Cove beat my record
sheepscreek 3 hours ago
So are they doing away with the Sol/Terra/Luna split already?
saaaaaam 9 hours ago
Pelicans please
atemerev 9 hours ago
Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
simonw 4 hours ago
Hah, this is a new one: first time there's been a complaint about the pelican before I've even posted one!
(I don't have access yet.)
saaaaaam 8 hours ago
Well you’re just no fun are you?!
maipen 9 hours ago
Very well said. It kinda describes how unrealistic these expectations are.
Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.
Very very unrealistic and wasteful.
droidjj 8 hours ago
saaaaaam 8 hours ago
ianm218 8 hours ago
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
alpineman 8 hours ago
That Astra ‘city scene’ is about as creative as Doha in real life (not very)
elzbardico 2 hours ago
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
cromka 7 hours ago
Surprised they haven't reset Codex usage on this occasion.
damsta 7 hours ago
I'd say it's because it's not available yet on subs
prometheus1992 9 hours ago
this is crazy! can't wait for the 27B distilled version of this.
E-Reverance 8 hours ago
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
wahnfrieden 10 hours ago
They're just announcing later availability. No launch.
paxys 9 hours ago
Every frontier release nowadays is "we've launched*"
* for a special group of customers that you're not in. Keep waiting peasant.
meowface 9 hours ago
That didn't happen with Fable 5.1 two days ago.
kegs_ 9 hours ago
pixl97 9 hours ago
I mean tell Nvida to 100x their hardware output and you'll get what you want.
iAMkenough 9 hours ago
Their announcement about later availability is unavailable to me now (500 error).
Great first impression.
semiquaver 9 hours ago
Guessing this one will never show up in cursor…
firemelt 9 hours ago
damn seems I should hold off my claude subs
alex7o 8 hours ago
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
jiraiyasarutobi 9 hours ago
It saturated most benchmarks. WTH
rbreve 8 hours ago
Where is the cure for cancer?
XCSme 8 hours ago
We need CancerBench
azan_ 7 hours ago
Didn’t Moderna use AI for development of their melanoma vaccine (which has recently shown spectacular results)?
dakolli 5 hours ago
Lmao, come on dude, anyone whos used these tools for research knows it makes them lazier, less interested and dumber. You really want disease researchers become sloppers too?
Rover222 8 hours ago
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
retired 8 hours ago
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
yodsanklai 6 hours ago
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
Obluness 7 hours ago
That seems promising ?
tinyhouse 9 hours ago
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
balefulboy 9 hours ago
72 to 74 on DeepSWE is AGI
mrcwinn 5 hours ago
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
jonplackett 9 hours ago
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
dowakin 9 hours ago
So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
damsta 8 hours ago
Why release it now instead waiting those few days until it is available for everybody?
balefulboy 8 hours ago
Because they saw how much hype Glasswing was getting in April
damsta 7 hours ago
From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
m3kw9 2 hours ago
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
dopa42365 8 hours ago
like eh 2 days ago it was the usual "too powerful to release"
https://www.reuters.com/business/openai-says-upcoming-model-...
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
guilhermeasper 10 hours ago
That was a quick pull out.
ChaseRensberger 6 hours ago
when do i get to go to the moon
johnnyApplePRNG 9 hours ago
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
camillomiller 6 hours ago
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
perching_aix 6 hours ago
gpt-6-astra-ultraspeed when?
Readerium 9 hours ago
dang 9 hours ago
Link added to toptext. Thanks!
HardCodedBias 7 hours ago
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
bbor 8 hours ago
To be, or not to be, that is the question:
Whether 'tis nobler in the mind to suffer
The slings and arrows of outrageous fortune,
Or to take arms against a sea of troubles
And by opposing end them. To die—to sleep,
No more; and by a sleep to say we end
The heart-ache and the thousand natural shocks
That flesh is heir to: 'tis a consummation
Devoutly to be wish'd.
...
And thus the native hue of resolution
Is sicklied o'er with the pale cast of thought,
And enterprises of great pith and moment
With this regard their currents turn awry
And lose the name of action.brcmthrowaway 8 hours ago
Anthropic in tears today.
nullbio 22 minutes ago
Anthropic don't care, they don't want their products to be used by general audiences in any serious manner. Their interest is in selling to megacorps and using the models for themselves internally to swallow industry, and drumming up AI fear to regulatory capture to shut down the businesses that do want to make AI accessible to the people.
colesantiago 9 hours ago
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
jdee 8 hours ago
'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
NichoPaolucci 8 hours ago
Yes. I had Codex rewrite and fix all of this in one shot earlier today (using Typescript). Unfortunately, I can not show you the code, because I do not know how this "git" program works but the AI keeps talking about it.
colesantiago 8 hours ago
Yes.
This is just another problem for the AI Labs to solve.
amazingamazing 9 hours ago
We have such great AI and cannot keep a static site up?
torginus 9 hours ago
Yeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.
pixl97 9 hours ago
Sometimes being the busiest site in the world for a few moments is difficult.
amazingamazing 9 hours ago
Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
pixl97 9 hours ago
gchamonlive 9 hours ago
That's the scientific positivism fallacy exemplified in one question.
gorgmah 9 hours ago
Yeah, apparently
agumonkey 9 hours ago
still hugged
HSO 7 hours ago
people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)
the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing
agi deus ex machina descending from the icloud ftw!!!
pathetic :)))
frozenseven 10 hours ago
Release the Kraken!
Pym 10 hours ago
I saw it
karim79 6 hours ago
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
tonyhart7 8 hours ago
its insane how they are dropping this after fable
Pieczasz 9 hours ago
Oh brotha, here we go again, it's so over again, as every week nowadays
wieiw1 9 hours ago
I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.
But thank you for spending other peoples money to give us the tech regardless!
dearing 8 hours ago
no results
Onavo 9 hours ago
The jump in scientific performance is non trivial.
dakolli 6 hours ago
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
nullbio 17 minutes ago
That's a policy and distribution problem, not an AI problem. Anthropic is doing their best to make it a reality though. They'd love nothing more than to shut down distribution and become the sole gatekeeper of everything AI.
ChrisGammell 8 hours ago
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
jpatten 7 hours ago
Yeah I was really excited to see the KiCAD example. Curious how useful it is in practice.
ChrisGammell 5 hours ago
I can't help thinking "doesn't matter much unless it's perfect" because if someone is using this to build a board (cool) but then it's not flawless, troubleshooting will be quite tough as a novice. Like, say, when I start digging into the web code generated by a coding agent.
I am most excited about it bringing down the barrier so more people join in on hardware fun, so hopefully it will unlock folks that stayed away in the past.
kingjimmy 8 hours ago
bro wtf is this website and why does it take 500mb of memory... smh.
danieltk76 8 hours ago
great, but nobody can use it for another 100 days right?
SneakyZero 6 hours ago
"GPT‐6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business"
holoduke 7 hours ago
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
bdangubic 8 hours ago
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
Brainspackle 10 hours ago
huh?
Maxforever 2 hours ago
Mhm
ealready_value 10 hours ago
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
paxys 9 hours ago
Why is this flagged ?
dang 9 hours ago
The link was 404ing quite a bit and several previous submissions got flagged as well.
consumer451 9 hours ago
It's still down for me, in the EU.
rvz 9 hours ago
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
Supermancho 9 hours ago
> Did humans deploy the model, Or did the model deploy itself?
> It sounds like "AGI" just stands for "IPO" as it always has been.
People don't usually respond to noise.
rvz 8 hours ago
Here's an idea, maybe answer the question before responding since you saw it?
What do you think?
adan1719 7 hours ago
AI releases are like religious ceremonies. You are not allowed to disrupt them. The new system card is the gospel.
bicx 10 hours ago
Dead link for me
jonplackett 9 hours ago
The launch video is incredibly cringe.
BadBrands 6 hours ago
So they’re copying Gemini with the whole star motif?
I guess it makes sense they are unoriginal.
like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas
steve1977 27 minutes ago
"distilling"... ;)