GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price (openai.com)
667 points by crorella 5 hours ago
proxysna 3 hours ago
I am yet to spend $200 on deepseek this year. Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500. Deepseek is faster, IMO intelligence difference is negligible and it so much cheaper that i no longer care about how much i use it. I never hit any daily/weekly quota or anything like that while working or tinkering. At this point i am OK with being 6 months behind the "frontier", purely on bang-for-buck basis and who cares which shadowy government gets my data.
rapind 2 hours ago
So I took Deepseek V4.1 Flash for a spin maybe 2 weeks ago now (before Luna 6 and Sol 6 were announced), and I racked up $100+ in about 2-3 days. It was pretty great, but it uses way more tokens (TPS is fast, but it's way more tokens per turn) than Sol 5.6 which I found to be about it's equivalent at the time (on medium or high, with DS on max). My cache rate was around 98-99%.
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
jbellis 2 hours ago
There's a bunch of skepticism in the replies but I ran over 100 tasks against DeepSeek 4.1 Flash and Sol (among others) and I can confirm, it is in fact a little smarter than Sol and a little more expensive than Luna. https://slopcop.com/power-ranking?pricing=api
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
cpursley an hour ago
taylorfinley 2 hours ago
Were you using OpenRouter? I've used 1.8bn tokens in the past week from DeepSeek themselves and 99.2% were cache hits. Total cost was $18.13 usd.
faitswulff 2 hours ago
rapind 31 minutes ago
InsideOutSanta 44 minutes ago
Same experience. I often see people say how little they spend on DeepSeek v4.1 flash, but when I put 60 bucks into my account, it was gone in a few days of non-exclusive use. I'm actually curious what the difference is. I used it through pi and opencode, but the harness seemed to have no obvious impact on usage.
christophilus 2 hours ago
This is my experience, too. It's a great model, but it burns tokens if you use it heavily for work on complex domains.
Edit: others have noted the provider and harness matters. My experience is with opencode.
MisterMunchkin an hour ago
How have you spent hundreds of dollars? I’ve only spent 11 and I’ve been using it for four months!
pimeys 2 hours ago
How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.
What harness you are using?
rapind an hour ago
the__alchemist 2 hours ago
guluarte 2 hours ago
same, used a wrapper around cc and i was spending up to $30 a day with basic stuff
taylorfinley 2 hours ago
sillysaurusx 3 hours ago
It’s easy to hit your quota. “Speed up the compilation time of this C++ codebase. Feel free to use several subagents to search through the files in parallel.” That’ll cost you about $200 for a codebase of ~1,000 files.
Subagents are like trading derivatives. You can lose as much as you want.
abixb 3 hours ago
What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
MintsJohn 3 hours ago
KeplerBoy 2 hours ago
ldng 2 hours ago
kruipen an hour ago
zer00eyz 3 hours ago
catigula 3 hours ago
gbacon 2 hours ago
> Subagents are like trading derivatives. You can lose as much as you want.
Excellent pithy warning.
sheepscreek an hour ago
Sadly this is true - for individual folks on the lower end of the spend spectrum.
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
proxysna 3 hours ago
Afaik there is just pay-as-use with Deepseek
deadbabe 3 hours ago
Why use subagents at all
FearNotDaniel 3 hours ago
nater5000 3 hours ago
giancarlostoro 2 hours ago
qarl 3 hours ago
HDThoreaun an hour ago
phyalow 3 hours ago
pnw 2 hours ago
I spent the weekend trying Deepseek 4 Pro on a Linux porting project and it led me down a complete rabbit hole where Linux wouldn't even boot by the end of the weekend. Waste of $120. Switched back to GPT 6 on Monday and Linux is booting again and I'm making progress.
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
forsalebypwner 42 minutes ago
> Deepseek 4 Pro
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
rapind 27 minutes ago
That's an unfortunate experience. Think of v4.1 flash as actually v5.0 flash. It's night and day compared to the 4.0 flash (and 4.0 flash was unintuitively better than 4.0 pro). I would re-evaluate with v4.1 flash. I'm not saying it better than Sol or anything, but it's in the ballpark.
pimeys 2 hours ago
You mean 4.1 Flash which is the first great Deepseek? The one that actually surpasses Opus in my books now.
jrflo 3 hours ago
There's a difference between "write this function for me" coding agents and "build this prototype from end-to-end". If you're doing the former, deepseek is fine. If you're doing the latter, it's not gonna work, and that's where the extra intelligence is most valuable.
throooooo 2 hours ago
This was my experience 3 months ago. I had an Android app that interacted with a Bluetooth device that I wanted to reverse engineer and build my own Linux app for it. DeepSeek was struggling really hard. Claude did it end to end after 3 or 4 prompts. To be fair, I was using a web interface for DeepSeek and the CLI for Claude; maybe that makes a large difference.
sreekanth850 3 hours ago
build this prototype from end-to-end, is this how people build serious software with AI?
jrflo 2 hours ago
thangalin an hour ago
mitthrowaway2 an hour ago
poilcn 2 hours ago
outside1234 3 hours ago
piterrro 3 hours ago
Oh buddy you have no idea what a good plan and agent harness can do with deepseek…
proxysna 2 hours ago
I am doing mostly hardware drivers recently, it works fine for complex work.
logicchains 2 hours ago
"build this prototype from end-to-end" works fine with DeekSeek V4.1 Flash, the problem occurs if you're not only building a prototype but want a finished product.
apitman 3 hours ago
I've been trying to use DeepSeek V4.1 Flash more and been very impressed. My current (very rough) rule of thumb is that an Artificial Analysis score of ~40 is the crossover point for "good enough" for most of the things I need to do with coding agents.
mchusma 3 hours ago
47 is my crossover for serious things (e.g. Grok 4.7 is below the line and GPT 6 Sol is above the line). I mean, Opus 5.5 is way better, but GPT 6 Sol still gets the job done for anything that doesn't require design thinking.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
apitman 3 hours ago
hmontazeri 3 hours ago
Had the same experience. I got rid of my pro sub of OpenAI. I’m really freaking impressed.
dzink 2 hours ago
nullbyte 3 hours ago
The intelligence difference between models like DS4.1 and Sol/Opus is NOT negligible.
bel8 3 hours ago
The premium price is only worth fo the hardest problems.
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
linuxftw 2 hours ago
zozbot234 an hour ago
In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.
tripleee 3 hours ago
If both DS4.1 and Opus can complete the tasks you throw at it at good enough quality the differences are negligible.
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
_benj 2 hours ago
nkjoep 3 hours ago
fragmede 2 hours ago
rmaxdev 3 hours ago
What do you do? I’m 200 bucks deepseek flash in about 2 months and it’s increasing
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
iammrpayments 3 hours ago
I put 5 dollars at deepseek a long time ago, and somehow it has never been fully spent, can’t imagine anyway to spend 200$ on that thing
proxysna 3 hours ago
Recently, drivers for a bunch of obscure hardware. Lots of c\c++, that i am ok with but not enough to make hardware drivers (i am just impatient). Just using pi agent with a few plugins.
marknutter 2 hours ago
thiht an hour ago
I've been using Claude Code at work and OpenCode for side projects for a few months. Every OpenCode model I've tried always felt subpar compared to Claude, but good enough. But it changed with DeepSeek 4.1 Flash, I've been using it for the past few days and I've come to forget I was not using Claude, it's a really good model and it's basically free for my usage (I used it almost all the weekend and spent ~$5)
holbrad 44 minutes ago
I think the only answer to this is you're just not using agents enough, because even with the very cheap pricing, it's still easy to rack up a large bill.
jwpapi an hour ago
A lot of people having different pricing experience. I think it’s important to understand that caching can differ, than if the agents spend waiting on code, or consume a lot of content. It depends on how you structure you codebase and how explorable it is, how much effort you set and probably some other issues.
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
LarsDu88 2 hours ago
I've spent $200+ on deepseek and this is for making a multiplayer FPS game. Trust me there are use-cases.
And no it did not deliver. A lot of it was re-done by Astra
abroszka33 13 minutes ago
> I've spent $200+ on deepseek and this is for making a multiplayer FPS game.
Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.
ApolloFortyNine an hour ago
The speed of deepseek is insane to experience after using claude code with opus for so long. Not only is the tps roughly 3x faster, but the round trip times are magnitudes faster.
mrbonner 38 minutes ago
$200/month is for Navier-Stoker grade problem.
zzleeper 3 hours ago
A bit tired of spending $200 out-of-pocket for openai. What do you use as harness? (for me the harness if half of the benefit... controlling my PC, working from phone, etc.)
simlevesque 3 hours ago
I use Claude Code + eternal terminal + tailscale + tmux + some custom skills to get notifications through nfty.sh.
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
IOT_Apprentice 3 hours ago
pimeys 2 hours ago
https://omp.sh/ has amazing defaults and it sips tokens. Works really well with DeepSeek V4.1 Flash.
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
kleinishere an hour ago
phyalow 3 hours ago
I have a server living in my home office, always on. I have a tmux session on it with vanilla Codex and Claude Code CLI, I can via my Ubiquiti network stack wiregaurd in to this box anywhere on the globe with just my laptop. Works super well for me. I also have some cheap shelley power plugs that I can use to cycle my PC’s power state if needed.
marknutter 2 hours ago
dolebirchwood 3 hours ago
OpenCode works nicely for me. You can connect it to the DeepSeek platform with an API key.
proxysna 3 hours ago
I use pi.dev, it is pretty minimal, but you can extend it however you want since agent has access to it's own documentation.
ThomasGlanzmann 3 hours ago
I use crush (https://github.com/charmbracelet/crush) with the following patches:
curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git amdzink 2 hours ago
Where do you do your inference?
giancarlostoro 3 hours ago
Are you just using it directly from them?
m3kw9 2 hours ago
Kind of ambiguous without saying token amounts and cost.
UltraSane an hour ago
Opus 5.5.is crazy good. I ask it to do things and it just writes the code to do it.
pdntspa 2 hours ago
I just ran a huge text/image extraction grudgematch against all the current inexpensive models except gpt-5.5/5.6/6 (due to some issues with openrouter and bugs in my code) and DS4 ranked very poorly. Accuracy winner was Gemini 3.8 flash with minimax M3 and qwen 3.8 placing, and the chinese models beat the incumbent (Gemini 2.5 Flash) on cost whilst keeping like 95% of the accuracy.
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
alfalfasprout 3 hours ago
It's trivial to hit that kind of quota if you're trying to execute on major projects. Especially as you start having dozens or hundreds of subagents investigating, prototyping, and working on different things.
FailMore 3 hours ago
API pricing?
revolvingthrow 4 hours ago
There was a model called Astra-Minor, found in the files a few days ago. I assume Sol 6.1 is this, as a last minute panic rename due to Sol 6 being underwhelming while Opus 5.5 turned out really strong. I can't really explain releasing Sol 6 in any other way, especially mere days ago.
nsingh2 3 hours ago
I don't understand why they didn't call 6-Sol just 6-Terra. It was 5.6-Terra level pricing with a perf jump.
r0b05 3 hours ago
I don't understand any of this naming man. It just gets more confusing.
Razengan 2 hours ago
outside1234 3 hours ago
jrflo 2 hours ago
Terra is the "missing middle" model and had no positive brand recognition. Sol was for intelligence, Luna was for efficiency. Luna max was cheaper and smarter than terra light and sol light was better than terra max.
hawk_ an hour ago
Because they asked the model what it wanted to be called?
swalsh 3 hours ago
I mean, when I upgraded my pipeline from terra to sol saying "it's the same price basically!" I was excited. Probably would not have felt as excited if it was just a version bump.
Not sure that's why they did it. But that was my experience.
ttul 3 hours ago
Was it a last-minute panic, or just OpenAI releasing an update when the had a bit more training under their belt to make 6.1-sol a whole lot better? Either way, I'm extremely pleased and will be giving this model a shot.
jumploops 2 hours ago
That seems likely, in the API GPT-6.1 Sol requires reasoning, just like Astra, whereas GPT-6 Sol (and Luna) allow "none"
jauntywundrkind 3 hours ago
Theo dropped their Sol 6.1 video and he says he can't tell you, but, shows enough to make it pretty clear. https://youtu.be/vu8X3YroB-w#t=5m30s
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
spwa4 3 hours ago
Sol 6 was pretty good the first 4 days or so of it's release. Then it was probably dialed back to a lower effort level and it became really bad.
minimaxir 5 hours ago
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
joshstrange 4 hours ago
> 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
redox99 4 hours ago
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
Aeolun 18 minutes ago
jorblumesea 2 hours ago
shimman 3 hours ago
onlyrealcuzzo 2 hours ago
If you're compacting every 5 minutes, you have a workflow problem - period.
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
ngruhn 2 hours ago
manmal an hour ago
Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.
AmazingTurtle 2 hours ago
you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.
apitman 3 hours ago
You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.
antonvs 3 hours ago
Try Gemini. It’s so cheap I often use my personal AI Pro account for corporate work, and most of the time it doesn’t matter.
ChickeNES 3 hours ago
codewithcheese 4 hours ago
you can config codex to compact at a higher context limit
_davide_ 3 hours ago
As a reference i burn 1% percent for every 40 minutes of sol on average
TuxSH 5 hours ago
Exactly half as expensive as Opus 5.5 in every API pricing metric
bigwheels 4 hours ago
And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
dotancohen 4 hours ago
beering 4 hours ago
TuxSH 4 hours ago
mmis1000 4 hours ago
sobiolite 4 hours ago
Infinity315 4 hours ago
jauntywundrkind 4 hours ago
dom96 4 hours ago
Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.
zeroonetwothree 3 hours ago
verdverm 4 hours ago
cache is typically 10%, is this OAI setting a new level at half, 5%?
crazylogger 4 hours ago
The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).
the_duke 5 hours ago
The GPT 6 release was ... not great.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
wkcheng 4 hours ago
I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
stldev 2 hours ago
My experience as well.
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
moshegramovsky 2 hours ago
100% hard agree.
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
pampas an hour ago
That's my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.
jsw97 an hour ago
After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
jrflo 2 hours ago
I'm in the same boat, I'll give 6.1 a shot but I'll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
bitexploder 3 hours ago
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
trentnix 3 hours ago
That's not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit.
r0l1 an hour ago
Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
chronogram an hour ago
Same here. Astra has been the best thing I've seen. Astra on Low has been my favourite thing so far. Higher levels just mean more cruft, not useful.
ozgung 3 hours ago
Maybe OpenAI was the only one pacing the frontier.
beebmam an hour ago
gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.
NorthSouthNorth 3 hours ago
I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
setnone 3 hours ago
yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it's just roll out pains
nxc18 4 hours ago
How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
user43928 4 hours ago
They never nerfed any model after release.
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
nxc18 3 hours ago
sigbottle 4 hours ago
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
moshegramovsky 2 hours ago
nxc18 3 hours ago
soulofmischief 2 hours ago
I have had the same exact experience. I feel like I'm working with 5.3 again. It is alarming how degraded the experience has become over the last month.
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
sunaookami 3 hours ago
gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: https://help.openai.com/en/articles/10306912-sharing-feedbac...
jeffybefffy519 an hour ago
Its almost like the "frontier" is a load of marketing bullshit and we should ignore it....
jstummbillig 4 hours ago
Eh. What? Is this common sentiment?
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
phoghed 4 hours ago
In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
copperx 4 hours ago
Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?
Marha01 4 hours ago
Eridrus 4 hours ago
Sol 6 definitely feels kind of dumb and worse than 5.6
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
nicce 4 hours ago
When GPT 6 Sol & Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can't trust it to do anything big alone anymore without babysitting.
the_duke 4 hours ago
On r/codex the sentiment seems to be quite wide-spread.
btbuildem 4 hours ago
That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
gradus_ad 5 hours ago
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
mixdup 5 hours ago
Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
luma 4 hours ago
Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
OliveronData 3 hours ago
chamomeal 43 minutes ago
john_strinlai 3 hours ago
digdugdirk 3 hours ago
trentnix 3 hours ago
interestpiqued 3 hours ago
dgellow 4 hours ago
CuriouslyC 4 hours ago
It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.
omalled 2 hours ago
sebzim4500 4 hours ago
Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
LPisGood 4 hours ago
Nvidia can start putting weights in silicon if model development slows down.
theturtletalks 4 hours ago
I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
serf 4 hours ago
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup 4 hours ago
colechristensen 4 hours ago
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
semiquaver 5 hours ago
What universe do you live in that you can look at the past six months and see anything like a plateau in capability?
Edit: removed a comment that was uncharitable and rude, for which I apologize.
arctic-true 4 hours ago
mixdup 4 hours ago
phoghed 4 hours ago
ActionHank 4 hours ago
xienze 4 hours ago
> sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead 4 hours ago
azan_ 4 hours ago
> Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
People were talking about plateau for years already.
eli 9 minutes ago
It would be weird if consumers were completely price insensitive.
djfjkfkffkkf 5 hours ago
China will do to llms what they did to german cars
bogrollben 4 hours ago
I guess I'm out of touch. What did china do to german cars?
CamperBob2 4 hours ago
Razengan 4 hours ago
ok123456 4 hours ago
We can only hope.
Razengan 5 hours ago
Why make a new account just to post this comment?
It's not even anything controversial..
simlevesque 5 hours ago
neta1337 2 hours ago
jorblumesea 5 hours ago
This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.
it's also why there have been so many calls for regulation and slowdowns.
LeBit 4 hours ago
Yup.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
nozzlegear 3 hours ago
0cf8612b2e1e 3 hours ago
There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
nojito 5 hours ago
Great for the consumer.
I remember when bandwidth was super expensive and now it’s dirt cheap.
vanviegen 5 hours ago
Not an AWS customer, I take it? :-)
iAMkenough 4 hours ago
That's relative to where you live.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
simianwords 4 hours ago
?! this model launch was around 10% of the dev day and the other time was spent on Dots and things other than models.
minimaxir 4 hours ago
That makes sense. There's not really much else you can say about it.
jimbob45 4 hours ago
Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
diego_sandoval 3 minutes ago
I find GPT 6 to be lacking in common sense when it comes to interpreting my prompts.
I have to be more literal with it than with GPT 5.x, otherwise, it sometimes does something totally different than what I want.
whatifitoldyou 2 hours ago
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
simonw 4 hours ago
I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv...
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
thefourthchime 22 minutes ago
I'm late with Pac-Man as well..
GPT 6.1 Sol — 91, ~9 min, $0.51 https://jonclegg.github.io/pacman-bakeoff/#gpt-6.1-sol
Opus 5.5 — 99, ~9 min, $2.00 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
GPT 6 Astra — 87, ~10 min, $2.42 https://jonclegg.github.io/pacman-bakeoff/#gpt-6-astra
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
Full gallery: https://jonclegg.github.io/pacman-bakeoff/
agar 3 hours ago
Did Medium not get a response, or is this a display issue?
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
UnboundedContex an hour ago
I always notice this too. Getting it right seems (psychologically for me anyway) to be a big part of "a good pelican" whenever I look at these. But doesn't always seem to correlate with increasing intelligence of models (measured via benchmarks, experience with the model etc.).
It's not frontier pelican without the back leg behind the bike frame IMO.
simonw an hour ago
Sorry about that, markdown bug, now fixed.
dankben 2 hours ago
Let's be honest, they're all guessing when it comes to rendering order
pazimzadeh an hour ago
Can someone explain to me why on these benchmarks like these a higher effort level often has a lower score?
For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive
https://openai.com/index/introducing-gpt-6-1-sol/#deepswe
Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..
zamadatix an hour ago
More thought can cause the important info to leave the context or hallucinated info to be enshrined in the context and later acted upon, especially in long horizon benchmarks like DeepSWE.
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
Nevin1901 5 hours ago
I love free market competition. We're getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence
jeffybefffy519 an hour ago
Spotted the person who hasn't used these yet....
phpnode 5 hours ago
What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?
az226 4 hours ago
Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
Aboutplants 5 hours ago
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
sockaddr 4 hours ago
This is it.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
pythonaut_16 4 hours ago
Maybe process maturity too.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
scrollop 5 hours ago
Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
geeky4qwerty 4 hours ago
jokes on me, I pay for all the subscriptions.
toasty228 4 hours ago
Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one
copperx 4 hours ago
It feels like an updated 4.6. It's fantastic.
copperx 4 hours ago
> I'm not the only one
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
mckirk 5 hours ago
No, we're pacing ourselves to have the time to evaluate the impact each new model could have, obviously.
sharpshadow 5 hours ago
Response to DeepSeek’s technical paper and competition.
LPisGood 4 hours ago
Which paper are you referring to?
wg0 5 hours ago
What's that in summary?
Wheen 4 hours ago
MisterMunchkin an hour ago
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
jonatron 5 hours ago
Probably just the singularity, no big deal
orbital-decay 5 hours ago
Versions is marketing, snapshots/minor variations are easy and the number must go up. Release timing is another OAI's marketing tactic.
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
denysvitali 4 hours ago
They're pacing the frontier
blmarket 4 hours ago
and seems like they're claiming Sol/Opus are not frontier (and only Astra/Fable are)
jchw 5 hours ago
It is the only way to reduce prices while making it look like a good thing.
lxgr 5 hours ago
Wanting to have the newer model than the competitor, presumably.
dandellion 5 hours ago
The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
lxgr 4 hours ago
mynameisjonny_ 5 hours ago
The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out
motoboi 5 hours ago
New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.
jesse_dot_id 5 hours ago
No.
SwabbyNat74 5 hours ago
Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!
agluszak 5 hours ago
They're releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5
tjwebbnorfolk 5 hours ago
Competition
mattnewton 5 hours ago
Anthropic’s IPO?
esafak 4 hours ago
Productivity is increasing as models get smarter; we are ascending the singularity. I'm serious.
colpabar 5 hours ago
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
infamouscow 4 hours ago
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
system2 5 hours ago
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 4 hours ago
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
andybak 3 hours ago
copperx 4 hours ago
thraway3837 an hour ago
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
Aboutplants 4 hours ago
“OpenAI's new Pro 500 plan offers OpenAI's highest usage allowance and comes with access to its new "Ultrafast" feature — it also costs $500 per month.
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
TomGarden 4 hours ago
They're really (finally?) starting to behave like a company bleeding money.
Our VC-backed subscription days are numbered
glaslong 4 hours ago
Alas, I did enjoy burning investor money on my taxis, movies and tokens.
onlyrealcuzzo 3 hours ago
> Our VC-backed subscription days are numbered
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
m3kw9 4 hours ago
I'm ok with whatever price they give out given they are not a monopoly and have competition, the lock in is minimum for me. This means they have legit reasons to send us this price plan. I don't believe they would shoot themselves in the foot when there is cut throat competition (Claude/opensource) out there.
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
LeBit 4 hours ago
Let’s pray Chinese models are not banned.
Madmallard 3 hours ago
how could u even ban them? lol
girvo 2 hours ago
user43928 4 hours ago
Unfortunate that the Ultrafast is only available with the $500 subscription.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
cactusplant7374 3 hours ago
Ultrafast uses 6x the usage. They probably realize that people will complain if the plan limits are too low. In any case, the TCO of the newer chips is supposedly lower. Hopefully everyone is on ultrafast eventually.
torginus 4 hours ago
I think that's by design - they're going to IPO soon so if they can get a significant percentage of users to switch from the $200 to the $500, they can 2.5x projected revenue.
glub 4 hours ago
Yeah, that's not going to happen. They are more likely to lose a lot of customers, unless Anthropic does the same thing.
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
latentsea 4 hours ago
seizethecheese 2 hours ago
adonese 4 hours ago
Very risky to do so especially considering how well is opus 5.5.
scottLobster 4 hours ago
Computer0 33 minutes ago
I am skeptical that individuals on the $200 and $500 plans make up that meaningful of a portion of revenue.
moregrist 4 hours ago
This is pretty typical product positioning. You want to sell to both high-end and low-end users, so you offer products at a few price points. Then it turns out that that middle is a much better fit for most users. So you start making the middle a worse fit to push most of those users into the higher tiers.
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch 4 hours ago
The 200$ plan was appealing because you got 4x usage for 2x the price.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
jpadkins 3 hours ago
latentsea 4 hours ago
At $500 per month, it's cheaper to just buy GPUs and use local models.
tripleee 4 hours ago
Have you looked at the prices of GPUs lately?
latentsea 3 hours ago
honkycat 4 hours ago
Wow, canceling my sub. Lets see how Claude is doing these days.
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 4 hours ago
Well, here is a breaking-news for you: the 20x from Claude is not a 20x on the weekly usage, it's a 20x on the 5h usage, while the weekly usage is simply double the $100 plan...
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
MCArth 4 hours ago
the_duke 3 hours ago
diffuse_l 4 hours ago
enraged_camel 4 hours ago
cmrdporcupine 4 hours ago
spiderice 4 hours ago
> According to the company, existing subscribers will keep their current limits for a time, and will later receive a one-time credit to help them make the most of their new reduced allowances
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
honkycat an hour ago
surgical_fire 4 hours ago
OpenAI is deeply unprofitable, particularly on those pro plans.
The only way is for prices to go up. Way up.
mrtesthah 4 hours ago
It really does look like OpenAI is trying to gradually get rid of their subscription plans. Every week there is noticeably less usage available to them while each new model release boasts substantially cheaper API token pricing. If this continues then the two pricing models will eventually be at parity.
surgical_fire 2 hours ago
A_D_E_P_T 5 hours ago
Looking at the token prices, if this is half as good as 6-Astra for 3D model creation in Blender, it's going to be an absolute game changer.
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
CuriouslyC 4 hours ago
From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It's slightly worse at spatial detail but better at aesthetics and little touches.
A_D_E_P_T 4 hours ago
Interesting! Can you share an example?
CuriouslyC 3 hours ago
ekun 5 hours ago
How is it with animations?
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 5 hours ago
You need to use the Blender MCP. There is an official plugin for this now, so the third party one can be avoided.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
A_D_E_P_T 4 hours ago
I've only tried animating models in Astra-6, and I was quite impressed! It's rarely able to one-shot things perfectly, but it usually gets pretty close.
lukan 4 hours ago
Have you tried fable? (I did small experiements and was satisfied, but maybe there are reasons to switch?)
A_D_E_P_T 4 hours ago
No, because I always hit my Fable quota (Max 20x) in 12 hours on simpler tasks, and I'd hate to need to buy tokens at API pricing.
therealdrag0 4 hours ago
After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.
jdprgm 2 hours ago
6.0 Sol was literally a week ago... Basically continuous integration for model releases at this point.
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
alright2565 43 minutes ago
They do, when I was on the $20 plan, I got shown a Luna Reserve model which had its own dedicated quota.
bayesianbot 2 hours ago
Interesting idea, but at the same time it is just so cheap that you can just run it with API pricing. I sometimes do even if I have available usage that I'm going to cap so I'll save it for bigger models
poisonborz 2 hours ago
As many have surmised, this may have been a panic rename of Astra 6.1
machomaster 20 minutes ago
Astra 6.1 light, not the full-fledged version.
aabajian 3 hours ago
Opus 5.5 is on another level, especially when it comes to mathematics implementations. You can drop it a PhD-level physical simulation (for example, a contrast-injection simulation for angiography in my case), and it just...implements it. With full-on WebGL rendering in the browser, from scratch (or using an existing library, if you prefer).
fraywing 5 hours ago
> GPT‑6.1 Sol matches GPT‑6 Astra at roughly one-fifth of the cost
Astra is a pretty impressive model. Excited to try this.
gobdovan 4 hours ago
They have also cut allowances for subscriptions in half. So even in the best case scenario it's about 2.5 times cheaper for Codex users. They just seem to have matched Claude Sonnet 5.5 *API pricing*, but from what I see online, it seems Claude Code now has a much more generous subscription allowance.
Tadpole9181 3 hours ago
Only for the $100 subscription, correct?
gobdovan 2 hours ago
jumploops 2 hours ago
If the Terminal Bench 4.0 scores are to be believed[0] GPT-6.1 is an incredibly efficient model.
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
modeless 4 hours ago
GPT 6 Sol is obsolete after only one week! I am glad that they are not afraid to update the models more frequently. The Navier-Stokes thing revealed that it took them only a week or two to train a model more capable than Astra, and I want the pace of public releases to keep up with that.
pmdr 4 hours ago
I had it write some code the other day, boy was it awful-looking compared to 5.6. Worked perfectly, but ugly nonetheless.
samuelknight 4 hours ago
Sol 6 was a flop. Nobody would have cared if it was called Terra 6.
holbrad 33 minutes ago
It seems pretty clear that this is a much larger model than Sol 6, and you can see this in the much lower generation times. I think this is also the main explanation for the $200 plan being cut in terms of API usage.
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
TomGarden 4 hours ago
Impressive improvements, but GPT 6 Sol came out 7 days ago, and this one will behave differently. The panicked pace is becoming a liability, maybe they should have waited and released this as the 6.0 release
enraged_camel 4 hours ago
They are behind, hence the panic. On top of that, Sam has been trying to do another funding round, so he's desperate to make the company look good.
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
atonse 4 hours ago
People said Astra was a gut punch and that Anthropic was reeling. (Opus 5 was almost universally panned)
The best thing is that we benefit from these constant back and forth gut punches :)
jhonof 3 hours ago
JacobAsmuth 3 hours ago
nzoschke 3 hours ago
OpenAI is feeling really competitive again.
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
A few more thoughts here https://housecat.com/blog/introducing-housecat-agent
intenex 4 hours ago
They released GPT 6 Sol literally 6 days ago. We've accelerated to a weekly model release cadence. That seems like...a big deal.
zarzavat 4 hours ago
It's more like they released GPT 6 Sol too early because they were under pressure and now they are releasing the real version. You cannot do anything more than minor post-training in a week.
toasty228 4 hours ago
Implying they don't have like 3 or 4 "models" (different quants, post training, plain renaming) on the back burner at any point in time to do exactly that
ychnd 2 hours ago
I think this was supposed to be new Astra, but it came out shit. So instead of shelving it, they found a way to make it still seem like progress.
tripleee 3 hours ago
It's a competitive strategy with Anthropic, not necessarily releasing as soon as they're done
nater5000 4 hours ago
>That seems like...a big deal.
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
resters 31 minutes ago
Frontier models are being used to obtain training data from users. We burn tokens teaching OpenAI how to make a cheaper model that is almost as good. I think the new $500/month pricing strategy is a significant misstep by someone who has clearly not tried Gemini 3.8 Flash or Deepseek 4.1 Flash.
ylsilva 4 hours ago
For most sane people, OpenAI is the way to go... A lot of usage with very good models, but you know that Anthropic is laughing all the way to the bank with Opus 5.5 being "the best" model right now... There are a ton of people (and companies) that will just refuse to use anything else than the highest benchmarking model in existence.
xyzzy123 4 hours ago
Whats interesting right now is that those people also see opus 5.5 as being CHEAP because it's something like half the price of fable or astra.
user43928 an hour ago
It is at least 3x cheaper than Astra in the 20x subscription.
So yes, it is clearly cheap in comparison.
jfrbfbreudh 3 hours ago
I have personal Max subscriptions for both. Nothing OpenAI offers touches Opus 5.5.
Stevvo 3 hours ago
I'm on both $20 plans. Got more out of Anthropic these last few weeks. But the value proposition shifts constantly.
iamdelirium 5 hours ago
I wonder if releasing this soon sort of validates the rumor that Sol 6 was just the Terra model they bumped up and slashed the price.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
squidbeak 4 hours ago
Whether it was or wasn't, Terra's absence shows Sol has replaced it as the new middle model.
hyperpape 4 hours ago
If that were true, they’d have axed their margins.
glimshe 5 hours ago
This is great. But maybe part of the motivation is that 6-Sol wasn't as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.
xkcd-sucks an hour ago
> Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
codewithcheese 3 hours ago
sol-6 is terra-6. They figure that no one was using terra and they could bring the speed and cost saving of terra distilled on astra, but rebranded as the more popular sol.
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
JacobAsmuth 3 hours ago
6.1 Sol is slower than 6 Sol: https://artificialanalysis.ai/models/releases/gpt-6-1-sol
However, that doesn't say much. You can just run a smaller model at a larger batch size to get higher throughput but lower interactivity.
5555watch 2 hours ago
It's not that terrible then if that's the case. 5.6 Sol was superb in my eyes, and while I've tried a million things, there hasn't been anything that 6 Astra improved or did better than 5.6 Sol, while costing a ton of time and resources. Including research topics, where it should have excelled. So even if 6.1 Sol is no worse than 5.6 but cheaper and faster, the gutted Pro200 might still make sense.
mkaic 4 hours ago
I got a popup in my Codex just now saying "Try out 6.1 Sol!" and so I clicked the button to try it, and intriguingly, it set my model selector to "GPT-6 Astra Light" which makes me think 6.1 Sol may be in some way just a lighter/distilled version of Astra? defo interesting, not sure if I should read too much into it though. I see no option for directly selecting 6.1 Sol in my Codex Desktop UI.
solarkraft 2 hours ago
I wouldn't read too much into it. I've gotten this popup for a model I hadn't had access to yet before and it resulted in what you describe.
slekker 4 hours ago
Astra Light is the default option in the UI, so likely a bug
skerit 3 hours ago
So their original plan was to axe Terra, but then introduce an "Astra Light" model a week later? They had a nice lineup named for a whole 3 months, and they're already messing with it.
recursive 3 hours ago
Wait, so is coding not solved?
seaal 3 hours ago
Just got access in Codex, looking forward to trying it out. Opus 5.5 has blown me away with what it's capable of doing, hopefully 6.1 will actually be a worthwhile contender.
Excited to tryout Decisions API as well.
AnodicElegy 3 hours ago
Price/intelligence comparison with Opus 5.5 on Artificial Analysis:
https://artificialanalysis.ai/?models=gpt-5-6-luna-low%2Ccla...
According to this, at Max it's better and cheaper than 5.5 Medium, but worse than 5.5 High. At Medium, it's better and cheaper than 5.5 Low.
neosat 4 hours ago
Do these benchmarks have any meaning anymore? And do the announcements seem less exciting now? (Not taking anything away from the advances we are making but it seems more incremental now?) The reliable way to tell if you'll like a model is reliable collage/X reviews to gauge a model's capability and then trying it out to see if you like the style.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
CharlieDigital 4 hours ago
I fully believe that these models perform better in benchmarks versus their predecessors, but in real world usage inside of real, production codebases? They feel just as flawed as ever. I honestly have not seen any significant improvement in a few months. The last thing where I felt "wow" was `/fast` mode and Deepseek.
paskejl 3 hours ago
Wholeheartedly agree. Astra was some improvement over 5.6-sol in the sense that I'd "argue less" with it, but still frustrating and still sloppy. I'm starting to feel people are not honest about their experiences, they do very simple things or have very low standards. The biggest improvement i've seen from Astra so far is speed.
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
CharlieDigital 2 hours ago
dom96 4 hours ago
Surprisingly (or maybe not) it matches the performance of Astra on my benchmark[1], but is much cheaper. It is also head to head with Opus 5.5 on both the price and pass rate, but edges it out slightly.
jjcm 2 hours ago
Here's a comparison of a image->html flow for GPT 6.1 Sol vs Opus 5.5.
GPT 6.1 Sol: https://html.non.io/lcars-gpt-6.1-sol
Opus 5.5: https://html.non.io/lcars-opus-5.5
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
jjcm 2 hours ago
Original design is here btw: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...
BrokenCogs 4 hours ago
Hardly any comparisons to Opus 5.5, which means it's not great
vb-8448 5 hours ago
The real announcement is the ultra fast mode ... Astra at 300t/s is insane!
tandr 5 hours ago
I am afraid to ask for the price multiplier here. And it will burn your weekly allowance not in 1 day, but in 3 hours now? Or just one?
TuxSH 4 hours ago
8x for subscription, 6x for credits/enterprise: https://learn.chatgpt.com/docs/agent-configuration/speed
az226 4 hours ago
6x
objektif 4 hours ago
Where do you see this?
vb-8448 4 hours ago
i'm watching the keynote
moinism 4 hours ago
Ok, but we need 6.1 Luna soon. 6 feels worse than 5.6 in our agentic use case.
kenzic an hour ago
Is it really 1/5 of the price if most people who use it are also losing 1/2 of their credits?
Aboutplants 5 hours ago
So when does Anthropic answer? Tomorrow?
iosjunkie 4 hours ago
hopefully the answer doesn't include an increase in cost/decrease in usage.
Alifatisk 5 hours ago
You live in a ping-pong.
dzogchen 3 hours ago
I feel a little salty about the plan changes. I wanted to upgrade to the $200 plan a day after it was blocked. Now it only includes half the usage unless for those that got grandfathered into the x20 usage.
InsideOutSanta 3 hours ago
Grandfathered for a whole month. You're not missing much.
lukehandcool 3 hours ago
What happened to "we urge you to urge us to stop moving AI so fast"?
KingOfMyRoom 2 hours ago
The main issue I have is how they nerf their models and the quality difference between API users and their subscribers.
medler 2 hours ago
What is the quality difference between the API and subscriptions?
t-sauer 5 hours ago
Wasn't 6 released like last week? I can't keep up anymore.
algoth1 5 hours ago
It was so underwhelming that it didn't even make it to chatgpt chat interface
oh_no 4 hours ago
Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
nsingh2 5 hours ago
Yes but it was underwhelming, so they seem to have rushed 6.1 Sol out. Also Opus 5.5 may have spooked them too.
SirMaster 5 hours ago
Do you need to? Do you always keep up with all the version bumps on the software you use?
prodigycorp 5 hours ago
API price cuts were obvious once they made their announcement changing how usage is counted.
These moves all make sense when you take into account the enterprise market.
SirMaster 5 hours ago
Guys, are we slowing down yet?
scottyah 4 hours ago
Yes, obviously. They're both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.
arctic-true 4 hours ago
It’ll be interesting to see what happens to the economics of this business if we hit a wall on peak intelligence but keep finding cool ways to lower prices.
alvis 5 hours ago
Cache is priced at $0.1/M, 50% as sol 6 and sonnet 5.5.
thefounder 5 hours ago
They need to fix Astra first. My main issue is with GPT in general is that unless steered it goes into building AI “sloppiness”/machinery that is not “needed”.
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
hlynurd 5 hours ago
Weren't there headlines just yesterday that they weren't releasing this due to safety concerns?
pkulak 5 hours ago
That was 6.1 Astra. And I'm assuming it's being tabled because it still doesn't match Opus 5.5.
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
cmrdporcupine 5 hours ago
6 Sol was worse than 5.6 Sol from my own experiences. Far worse.
Will see if this remedies things.
pkulak 5 hours ago
tedsanders 5 hours ago
GPT-6.1 Astra is what those headlines referred to. This is GPT-6.1 Sol.
thejazzman 5 hours ago
"no homerS -- we're allowed to have one"
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
AmazingTurtle 2 hours ago
regarding the new pro max subscription btw:
it's 500$ for 25x the plus usage, thats pro (max).
this implies that the old 200$ 20x pro (more) is now more like 10x the usage of plus.
they are slashing our subscriptions in half and make it "but we're more efficient!"
user43928 an hour ago
Which they are, so another way to see it would be that they make the API/enterprise cheaper.
However, I don't know how future larger models such as the cancelled 6.1 Astra will be priced.
If the price stays high, this would indeed be quite bad for the $200 subscription..
nilslindemann 2 hours ago
Thanks, but I need the money for the next laptop.
cannonpalms an hour ago
OpenAI is cutting their subscription's token value in half. Half. I don't think this is enough to help them compete with Opus 5.5.
jrflo an hour ago
Just for the $200 tier, which was formerly the 20x weekly usage of the $20 tier. Now it's 10x the weekly usage, just like Anthropic's $200 tier.
nicce 5 hours ago
It is hard to trust these scores. GPP 6 Sol has been so bad for few days.
zerof1l an hour ago
Am I the only one who is struggling to keep up with all these GPT model versions and which one to use and when?
mekpro 5 hours ago
Why they are not even benchmark model against Anthropic or anybody ?
tensegrist 4 hours ago
noting that input:output:cached is 10:50:1 for astra and luna but 10:50:0.5 for sol. this doesn't mean a lot without "tokens per task" information but it's still interesting for there to be a "dip" like this instead of a monotonic change in one direction or the other
m4rtink 3 hours ago
Price wars, those always end well, especially if you are loosing billions per day!
dangoodmanUT 4 hours ago
I can appreciate they show that Opus 5.5 is objectively better intel. and cost wise on multiple benchmarks
demibabs 3 hours ago
Is OpenAI’s Cloudflare turnstile new? Never gotten it before.
Either way, a little ironic…
unsupp0rted 4 hours ago
Excited for Gemini 4 at equal or better coding and a fifth of the price of 6.1 Sol
oh_no 4 hours ago
3.8 flash is more expensive than sol 6.0, google uses a LOT of reasoning tokens
sanex 4 hours ago
If it's so good why is Dots, also released today, based on Astra?
Lapalux 5 hours ago
It's relentless, isn't it?
MisterMunchkin 4 hours ago
Still too expensive. Needs to come down to $1 or less to compete with china.
ChaseRensberger 4 hours ago
very surprised by the sentiment against GPT 6.0 Sol, I've been using it exclusively since release and it feels like a cheaper astra to me. admittedly I haven't tried any anthropic models in a while other than small tests since i can't use my anthropic subscription in other harnesses (like OpenAI has supported natively for a long time).
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
unsupp0rted 4 hours ago
This is the first time I've seen praise for GPT 6.0 Sol: it's widely disparaged on Reddit and here in the HN comments too. My own experience likewise shows 6.0 making loads of silly mistakes, both for things 5.6 Sol is good at and things 5.6 Luna Xhigh is good at.
ChaseRensberger 4 hours ago
well i could certainly be in the wrong; i'm just speaking from my personal and likely flawed experience but i feel like i've noticed silly mistakes in every (llm) model that has been released (and that i've sufficiently used) and it hasn't felt like 6 Sol was much of a regression from 6 Astra (more than reported in both model cards), both of which ive very extensively.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
zf00002 5 hours ago
I can blow through my weekly on astra in a few hours; hopefully this really is as good.
skybrian 5 hours ago
I get decent results telling it to use Luna subagents for implementing commits.
joduplessis 4 hours ago
Sol is so good, honestly - the sweet spot for me. I've only ever found it stumbles when you don't give enough direction. But for idea execution - Sol is the GOAT.
cmiles8 2 hours ago
The only real question that matters at this point is what do the economics look like for OpenAI. If they make good money on this with sane accounting principles then great. If this is just throwing more gasoline on the pile of burning cash to avoid losing more inference business then this bubble can’t pop soon enough.
andsoitis 2 hours ago
Sol, Astra.
Eventually: black hole.
gavin_gee 5 hours ago
the race to the bottom on models is well underway. huge IPO's only really make sense for DC/HW lockups, and going vertical.
EugeneOZ 2 hours ago
Can't trust a company which can halve a subscription any moment they want.
par 4 hours ago
Ok when can i get this in codex?
ghm2180 4 hours ago
> At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
hacker_88 2 hours ago
AGI is here
aaronbrethorst 5 hours ago
Shots fired, half the price of Opus 5.5.
barrenko 5 hours ago
A night of no sleep this one will be.
epolanski 4 hours ago
DeepSeek and GLM made it impossible for "sota" to price any way they want.
itzikkatz 4 hours ago
I’m a bit disappointed with Sol 6.1. I suspect they didn't show the benchmarks and test results because it would have been embarrassing to reveal that their flagship model can't compete with the capabilities of Sonnet 5.5. That said, I still think this model is useful for a great many things, but it looks like Anthropic has the upper hand this time.
slopinthebag 5 hours ago
$2/10 is pretty cheap for a frontier model...
cmrdporcupine 5 hours ago
Aka "we made an oopsie last week and released what should have been called GPT 6 Terra with the name GPT 6 Sol"
tultra 5 hours ago
Still not available to me
Alifatisk 4 hours ago
I guess they released this because GPT-6 Sol was underwhelming, they didn't even release it to ChatGPT. It was basically GPT-5.6 Terra for the price of Sol. However, who doesn't like price cuts? Astra for the fifth of the price? Wow, OpenAI have been quite generous recently, I still have not forgotten their 90% price cut with GPT-5.6 Luna, and now this? Astra was truly a milestone, and now they are offering similar "intelligence" for cheaper price. Incredible.
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
throwitaway222 5 hours ago
So yesterday we were consumed with how this was being delayed because of safety, yada yada.
Guess not?
nimonian 5 hours ago
I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don't think this release in particular contradicts it too much.
minimaxir 5 hours ago
That model was implied to be GPT 6.1 Astra, not Sol.
murbard2 5 hours ago
Astra 6.1, this is Sol 6.1
MetaverseClub 4 hours ago
Sam Assman needs money to buy a new super car or private jet?
spwa4 3 hours ago
I was wondering why GPT-6 astra has been performing so incredibly bad on codex for the last week. This seems to be a repeating pattern, to dial the settings on the current models to the idiot setting, and then release a new model about a week later.
varispeed 4 hours ago
I wonder if this is the reason GPT barely works now and some chats have become not accessible.
prometheus1992 4 hours ago
where's the pelican?
sergiotapia 4 hours ago
Typically how long does codex take to update with the right model metadata for the release of a new model?
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
IshKebab 5 hours ago
I wish they'd list the environmental cost. My employer has an unlimited AI budget so I don't care about using Astra if it's just more profit for OpenAI. I care more if it actually uses 5x more energy.
otterley 5 hours ago
Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.
lp92 5 hours ago
Did you care about this when it came to your other computing needs? What PC/laptop are ypu running and how efficient is that?
IshKebab 4 hours ago
Yes I do. I've got a spare desktop that isn't too efficient (probably ~100W idle but annoyingly I've lost my power meter) so I don't leave it on even though I would like to use it as a server.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
codehorses 4 hours ago
Likely orders of magnitude different, this is a weak whataboutism.
Bolwin 4 hours ago
Look at Neuralwatt. They report energy usage with every call as well as aggregate statistics.
sergiotapia 4 hours ago
I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar/battery/nuclear.
rs_rs_rs_rs_rs 5 hours ago
I don't understand the point of this, why just now when it comes to llms. Why wasn't anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.
Edit: for context, just Steam alone has ~200million monthly active users.
JDups 4 hours ago
Because people find video games fun, though I suppose there's some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
paulryanrogers 5 hours ago
Considering Nvidia's hard shift to crypto and now AI, I doubt videogames are even in the same ballpark.
How many DCs are devoted solely to gaming?
lp92 4 hours ago
rs_rs_rs_rs_rs 4 hours ago
lbrito 4 hours ago
If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.
rs_rs_rs_rs_rs 4 hours ago
IshKebab 4 hours ago
I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
rs_rs_rs_rs_rs 3 hours ago
jdw64 5 hours ago
GPT 6.0 Sol was so terrible—I wonder if 6.1 Sol will be good?
formvoltron 5 hours ago
didn't 6 sol just come out a couple weeks ago?
Alifatisk 5 hours ago
Other comments have already addressed this.
lynx97 4 hours ago
Came here only to check if the pelican spam has made it to the top again.
dcchambers 5 hours ago
GPT-6 Sol released a week ago. Shortest model life ever?
cmrdporcupine 5 hours ago
Taking GPT-6 "Sol" outside behind the shed and giving it a merciful end is about the best outcome possible.
Huge misstep releasing it.
tandr 5 hours ago
The misstep was naming it Sol - it was Terra-level all along.
algoth1 5 hours ago
Bro, I'm a visual thinker
cmrdporcupine 5 hours ago
gxs 3 hours ago
I have used both extensively and I don’t care what the bench marks say - Claude has been way, way better and most importantly, predictable
I can handle issues much better if they are predictable even if the model makes mistakes — much more frustrating when the model is erratic
I find codex wanders off road more often and fails to see the “bigger picture” (as much as LLMs can see the bigger picture at least)
And tbh when it was first released Astral felt even worse
I’m being forced to use it right now and at the end of the day I’m making do so it’s fine, but Claude makes for a smoother experience
dools 34 minutes ago
All of a sudden getting competitive on token pricing over the past couple of releases tells me they’re about to kill subscription pricing big time. The subsidised tokens aren’t going to survive the IPOs but if they can capture baseline dev tasks at a cost competitive with open weight models through Luna then capture the frontier token spend as well they could be pretty well placed. The Jarvis bros aren’t going to be able to afford their dashboards though.
sehw 5 hours ago
I stopped using LLMs. I shit you not. My life got better.
Starlevel004 5 hours ago
Okay, now price cut 6 Sol (and rename it to Terra again).
amelius 5 hours ago
If these models are so smart, can't _they_ select the right model for each task?
skulk 5 hours ago
the right model for the task is the one that transfers the maximum amount of USD from your pocket to the provider's bank account.
amelius 5 hours ago
No because then I'll go to the competition.
Mkengin 2 hours ago
Github is trying to do that: https://github.blog/ai-and-ml/github-copilot/project-hydrafu...
mholm 5 hours ago
Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.
condour75 5 hours ago
I guess the question is, does the Dunning Krueger effect apply to models? The dumb ones might think they're up to the task.
aleph_minus_one 5 hours ago
Why don't you simply ask the respective model which model is best for a specific task? :-)
amelius 5 hours ago
Because it is more work?
SkyBelow 4 hours ago
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
godwinson__4-8 5 hours ago
Let's all boycott and move to Claude until they release 6.1 Astra. I don't like to be teased.
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
ColonelPhantom 4 hours ago
Isn't Anthropic doing the same, with Opus 5.5 being out while Fable/Mythos is still on 5.1?
wren6991 4 hours ago
It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
godwinson__4-8 4 hours ago
Is this due to a similar safety concern or just because it's not ready yet for one (or more) of a myriad of possible reasons?
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.