GPT-5.6 Sol Pricing Cut by 50% (openrouter.ai)

466 points by Topfi 14 hours ago

pimeys an hour ago

The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between the models, and Kimi K3 was the first one that actually could compete or be even better than Opus or Sol in our use cases, with a fraction of the price. All our developers use K3 as their programming model, and it now powers a big part of our systems instead of Opus and GPT. Surprisingly the new Sol pricing is quite similar to K3...

Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.

And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.

gagan2020 31 minutes ago

I shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.

tkgally 37 minutes ago

I was struck by a video ad that Google released yesterday with testimonials by three developers about using Gemini 3.7 Flash [1]. The point they emphasize most is price, followed by latency. The marketing strategy definitely seems to be shifting.

[1] https://youtu.be/kacf2bib-X0

ywvcbk 42 minutes ago

> Opus or Sol in our use cases, with a fraction of the price.

I assume it's highly use case dependent, though?

Even before the price cut seems like Sol was price competitive with Kimi

https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhig...

And now it should be considerably cheaper

pimeys 2 minutes ago

Long-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than Opus right now.

You cannot just look, at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.

eru 33 minutes ago

> And these models are not going away, nor their prices going up [...]

Well, DeepSeek just raised prices.

pimeys 5 minutes ago

And Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.

netsec_burn 10 hours ago

After using Claude for a long time, I tested Sol 5.6 for the first time today. Love it, its an incredibly capable model and uses far fewer tokens/time thinking. Its what I imagine Fable would be if I haven't been downgraded on every conversation - even after completing the verification program. I think I may cancel my Claude subscription finally.

jchw 9 hours ago

I think Fable's dominance is overstated. It definitely has the lead, but quantifying what that lead actually is is really hard. I'm using GPT 5.6 Sol to do some shit that I personally would consider "crazy" - low level undocumented hardware driver alchemy, reverse engineering highly obfuscated code, even a bit of screwing around with a rendering engine in Vulkan, really just about the most complex tasks I can get any model to do, and it does great. For the more advanced stuff, it definitely needs the effort bumped. But even with the effort bumped, the token usage really doesn't seem to skyrocket too badly until at least you hit xhigh and max, which really only seem to be necessary if you are doing genuine crazy stuff, so it's not that bad. I did similar stuff with Fable. In fact, I went directly from an Anthropic subscription with Fable to an OpenAI subscription with Sol, more or less, and it really felt pretty seamless. If anything, I was thrilled to realize how much I actually preferred Codex CLI, to the point where I started using it at work too.

Fable seems to be generally more impressive at outputting one-shot web apps. I'm not really saying that to try to downplay what Fable can do, it's just that if I compare the two, this is one of the few definitely noticeable areas that you can easily demonstrate. Obviously, one-shotting programs is much better as a demonstration of a model's capabilities than it is practically useful (not that it is useless, but hopefully my point is understood).

However, whatever Fable truly is better at, one thing I really like about GPT 5.6 Sol is even harder to quantify: taste. GPT 5.6 Sol outputs are still LLM outputs and they contain many things that people would probably consider "Claude-isms" for better or worse, but overall I really prefer the GPT 5.6 Sol output. I find it to be generally more tasteful. Hard to quantify, but when talking to people I've had enough people seemingly agree with me to convince me that it really is true.

andrewingram 9 hours ago

I used Sol to extract the remaining decryption keys from the Super Mario Maker 2 (Switch) game files. Someone had previously extracted all the keys from the original release, but not any of the new ones from updates. Not only did it succeed, but it helped me understand the data sufficiently to add support for “Super World” rendering to my level viewer (which I made back in 2021), eg the little widget at the top of https://www.smm2-viewer.com/players/B16-306-GVG

I was very pleasantly surprised to find Sol wasn’t obstructive over what was clearly a very grey area endeavour.

mjhagen 4 hours ago

alchemist1e9 7 hours ago

eli_gottlieb 6 hours ago

hatthew 7 hours ago

I feel like those examples are considered difficult because they're niche topics, but aren't actually all that difficult in a general sense. What I consider truly difficult are things like taking a ticket and implementing it in a preexisting codebase, using a clean and reasonable design that fits the existing style and makes sense to a human, and avoids the footguns I learned by working with the codebase for over a day.

jchw 5 hours ago

user43928 4 hours ago

edg5000 8 hours ago

FYI I run it consistently in xhigh regardless of difficulty of the task at hand. I remember high being very fast, but I'd rather wait a bit more and get better output. AIs are insanely fast compared to me anyway, even on xhigh. Consumes more usage, but even at 100 EUR/m I don't hit limits.

saghm 8 hours ago

jchw 8 hours ago

FooBarWidget 4 hours ago

cgannett 8 hours ago

And Mai-Code-1.1-Flash seems like a really good cooperative player to GPT 5.6 Sol. You get Sol to help you make a detailed plan, and Mai codes it up and you can get pretty decent code out the other end without too many tokens if you are careful.

swingboy an hour ago

manmal 6 hours ago

wonnage 4 hours ago

AI-pilled obsession with "taste" is bordering on insanity

It's just vibes

mfru 4 hours ago

calvinmorrison 9 hours ago

things that are alchemical are rarely alchemy. That is to say things are very fiddly but stick a room of monkeys on typewriters, a schizophrenic developer with HolyC and adderall or an LLM, persistence is the key to many of these things like drivers, extracting keys from vintage security domains, etc. Dropping into xdd to a human is a chore, not for an LLM.

jchw 8 hours ago

kyxsc 8 hours ago

Sol is way too eager to hone in on small details and ends up with massive over-engineering. Fable does it too - to be fair - but noticeably less.

After extensively using both on Max 20x plans, I've concluded that Fable is better for problem solving and coding, whereas Sol 5.6 Ultra shines in debugging specific issues: tackle a problem with Fable then leverage Sol to clean up, double check, or fix specific issues.

Fable (imo) had the edge on the $200 plan, but after this 50% reduction I'd say Codex is better value by far and there's no contest.

---

Using Fable as the orchestrator and delegating tasks to Sol 5.6 Ultra via the codex plugin in Claude Code yielded good results, but still there was a lot more over-engineering (thus time and tokens spent) than Fable by itself would've done.

Both models suffer from doing-too-much. But both models are fundamentally really smart and knowledgeable. I think it's really close and pricing cuts really spice things up for us consumers! Sol is a clear winner in the value department and the $100 plan is enticing!

---

*Claude Code usage is reducing by 33% in 2 days, Wednesday August 19... cmon anthropic: clau.de/cc-50-promo

andreygubarev 3 hours ago

Yeah, I very much agree on this. I think Sol and Fable code quality is on par. Maybe Fable is just a tiny bit better, but Sol compensates with its ability to work through things, while Fable, in my experience, generally tends to avoid solving problems that require many LOC.

However, I think these are very different models in terms of orchestration. Long-horizon tasks are way more predictable with Fable. It just doesn't lose track of details. Thus I ended up building a small wrapper around Pi (where I run Sol) so that CC can delegate via background tasks, automatically wait for completion, and do what was one of the most effective parts - steer Sol toward simplicity, getting Sol out of code-review infinite loops (Pi calls for Codex review to ship better, but generally gets stuck on P2 and results in vastly overengineered work).

One of the worst experiments was enforcing coverage at 100%. Only Sol, with an enormous amount of code and significant pushback (on architecture decisions) to Fable, was able to reach it. It made me think this is somehow related to overengineering in general, so that instructions on acceptance criteria in claude.md plus proper DX (e.g., Lefthook) actually led to okay results. It mostly helped that responsibilities were clearly split: Fable designs architecture, Sol handles coding and debugging.

greenavocado 8 hours ago

Sol w/ Effort -> Low

kyxsc 8 hours ago

iJohnDoe 8 hours ago

Interesting the use of Max and Ultra. I don’t doubt the complexity, but would someone use Max or Ultra on Typescript or Go, for example?

Is it more about just avoiding any mistakes? Seems like that would be costly when medium or high would work fine?

kyxsc 8 hours ago

leokennis 4 hours ago

5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more/less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.

FL410 9 hours ago

Fable feels less cumbersome to work with, but it is SO DAMN ANNOYING with the refusals that I'm leaning more and more on Sol, and very much looking forward to GPT6. Just seems like Anthropic is trying their hardest to ruin their reputation and user experience.

paxys 9 hours ago

Remember, Dario knows best

teaearlgraycold 8 hours ago

I’ve never had it refuse anything. Even vulnerability searching in my codebase.

ChadNauseam 4 hours ago

matheusmoreira 5 hours ago

I too switched to OpenAI after I got sick of Anthropic's constant "safety" downgrades. Sol is definitely a breath of fresh air.

> even after completing the verification program

Was it easy to complete it?

I ended up in some weird state where I can't even attempt the verification at all. Opened the Persona tab once, closed it and then it never opened ever again. It says a verification precheck failed.

Even without TAC, Sol doesn't seem to get blocked very often. Fable would downgrade to Opus if I looked at it wrong.

ec109685 9 hours ago

Fable is still the best there is. Sol close second but I find it gets way to stuck on details.

Also Opus 5 is fine if your codebase is simple.

curreylabs 9 hours ago

And you arent writing English

boorang 4 hours ago

I was pleasantly surprised to find that the GPT models are much stricter in adhering to my AGENTS.md guidelines and heuristics than Claude.

Aargau 9 hours ago

I'm also in Anthropic Cyber Verification Program, but they specifically exclude Fable, just goes up to Opus 5.

I hear you on the downgrades, I'm 13/13 on downgrades, and last downgraded me to Sonnet for asking for reasoning chain.

shibaprasadb 2 hours ago

I cancelled my subscription recently and moved to Sol. So far - it has been a great experience. The only aspect where Fable/Claude is better I feel is doing some research from the web and summarising the facts.

nutribeatApp 32 minutes ago

I generally prefer Fable but in my experience Sol is a much better web researcher

arcrs 39 minutes ago

output token efficiency bruv. u never go wrong with it

znnajdla 5 hours ago

Sol is my daily driver but there are still times I reach for Fable when Sol doesn’t cut it. Just yesterday for example, I was trying to build a self-modifying hot-reloaded agent harness in Elixir for fun and Sol just kept doing silly things like thin wrappers and unnecessary abstractions. Fable handled the task elegantly. Sol is really good as a reviewer for finding bugs due to its thoroughness however.

impulser_ 2 hours ago

It's the complete opposite for me. The model might be the worst model I have ever used when compared to other models in the class. You just can't get it not to just write the most enterprise over complex over engineered solutions for every little thing you ask it to do.

It the first model to actually make me pissed off to use AI. I absolutely hate the model so much.

I don't even want to see the codebases this model is fucking up.

It might just be good at finding bugs that about it. That all I would ever use it for just because it works harder than Claude models.

nurettin 34 minutes ago

Used claude since 7/2025. Switched to codex after fable got blocked. It was still 5.5 but I knew they had to come up with something. As soon as I switched, wow. It wasn't super intelligent, but it was stable. Every day it was the same performance. This consistency is definitely worth paying for.

vdntp 5 hours ago

you should check out the codex desktop app. people who've been using claude code for a long time will surely be surprised.

tamimio 9 hours ago

Yeah at this point claude is overrated, overly expensive, weird writing style (elliptical), and the worst part is the aggressive guardrails that even normal convos get interrupted, meanwhile openAI is still I would say at the normal balance, if you ask something too obvious or direct it will stop you other than that, it work flawlessly, plus, I have yet to hit the limit despite heavily using it these past weeks.

ChadNauseam 4 hours ago

> if you ask something too obvious or direct it will stop you other than that, it work flawlessly

What on earth are you asking it?

jiaosdjf 2 hours ago

The final straw for Claude was its refusal to give me a list of the most recent rapes reported by the BBC and basic information about them (location, date, names, just things reported in mainstream media). It outright REFUSED to complete this task.

I will not be told what I can and can't do by AI and I will no longer be supporting American companies run by despicable people. GPT only gets my money right now because its so fast and cheap but I'll be back to Chinese models in no time.

saidnooneever 4 hours ago

sol is much better imho than Fable but i can understand if they will perform wildly different for different people with different levels of expertise aswell as different needs. I dislike fable myself it doesnt really work for me.

Sol also doesnt _really_ work but it sort of tricks me into thinking it does more convincingly :p.

cancelled my subscriptions few days ago. (was on 100$ ones, not sure if there is diff in quality for higher tiers or not.. there might be that too).

what i hate the most is that they will make any obvious mistake you do not tell them to avoid. then on the next plan to fix it, your token limit is hit at step 4/5 -_-. Both models seem incredibly good at that mostly...

for tasks outside of coding and program design i do find them quite useful. like devops crap. maybe because i hate that, i like their help there more.

yieldcrv 5 hours ago

Claude as a harness at all really spends too much time before giving user feedback

Its a crutch that is no longer competitive

petesergeant 5 hours ago

I have a "strategy / life-coach" project, and was surprised at how much better Sol is than Fable on it, as I've found Fable to have the edge for most things for me so far. But Sol: questions were better, insight was better, it got the brief better.

Kye 8 hours ago

Sol has held stuff for a while to do the same sort of hazard checks I assume Fable is doing, but it always releases them. I think that's the better way to handle it rather than preventing me from seeing how far I can get generating schematics to use in Minecraft. Currently: a mostly normal voxel house.

Razengan 10 hours ago

I recently tried Claude again after several months, to see if it was any better at something Codex has been struggling with…

They STILL don't have an option to "Sign in with Apple" on the website, but they do for Google??!? (and on iPhone of course)

Screw that asinine UX

(and no it wasn't better than Codex at this particular task)

nozzlegear 6 hours ago

That was an issue at least a year ago. I had signed up for a claude account on my iPhone and then wanted to sign in on my laptop but nope, not possible. Insane they still haven't fixed it.

Can somebody at Anthropic tag claude in slack or whatever goofy shit you do and ask it to add Apple OAuth to your website? Clearly humans aren't testing it.

Razengan 41 minutes ago

resonious 4 hours ago

Where is the official source for this?

OpenAI's docs still show non-discounted pricing https://developers.openai.com/api/docs/models/gpt-5.6-sol

mirzap 4 hours ago

It's discounted if used via OpenRouter, not the official API.

londons_explore 2 hours ago

Doesn't seem like a smart business move.

You're literally encouraging someone else to come in and steal your customer base,

dannyw 5 minutes ago

rafaelmn 24 minutes ago

ywvcbk 37 minutes ago

podgorniy 2 hours ago

krzyk an hour ago

So the title is misleading, intentionally.

oblio 3 hours ago

That's weird.

JCharante 2 hours ago

not-kinsale-joe 3 hours ago

onlyrealcuzzo 7 hours ago

This sure looks like a race to the bottom to me, and I love it.

If Sol isn't the best model, it is up there...

You don't cut the price of the best model for no reason...

dvt 4 hours ago

> This sure looks like a race to the bottom

Always has been. My prediction is that both OpenAI and Claude will go bust unless they deliver a killer product. And unlike scrappy startups, they have a pretty serious deadline because creditors will come a-knockin'.

There's little to no functional difference between Kimi, Qwen, Sol, Opus, etc. All flagship models are within like 1-5% of each other and the real moat will be what's always been the hard part: making a good product.

matheusmoreira 4 hours ago

> All flagship models are within like 1-5% of each other

Don't know about that.

I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc.

Ran a complete Fable/max code review. Took over a month on a subscription. Now I've switched to OpenAI and am repeating the exact same review with Sol/max.

It's still not done yet but preliminary findings suggest Sol can only reproduce 70-90% of Fable's findings. So I think these models aren't as close as we've been led to believe.

beacon294 2 hours ago

roncesvalles 2 hours ago

The problem is that most of the volume doesn't come from proprietary products, it comes from API use which has no stickiness.

Claude already has a killer product (claude.ai/chat is a Swiss army knife) but just relying on people typing stuff into chat is not enough to sustain the company.

The other strategy is entrenching yourself as the LLM of choice into existing products (like ChatGPT is on Apple products).

sumedh an hour ago

> All flagship models are within like 1-5% of each other

Depends on your use case. the Chinese models are not there yet.

blfr 2 hours ago

There is a massive difference even between Opus and Fable, same provider, before various harnesses and other optimizations come into play. Don't be deceived by rankings and benchmarks, try for yourself.

twelvechairs 4 hours ago

If you believe https://artificialanalysis.ai/

This is basically undercutting KimiK3 and Grok 4.6 where previously utilised gad soke advantages but was a step more expensive

andai 7 hours ago

You do it if you can afford to do it and your competitor can't.

psadri 7 hours ago

Who are OpenRouter’s competitors?

aurareturn 5 hours ago

fileeditview 5 hours ago

That kind of is a reason.

livinglist 4 hours ago

The lower the merrier.

oblio 3 hours ago

Competition is good for users.

livinglist 3 hours ago

Fergusonb 11 hours ago

Luna saw a huge jump after the price cut and is one of the more competitive models at the new price on openrouter.

Maybe they want to see how much market they can grab with Sol?

This might help but there are already cheaper models with Sol's intelligence more or less, the most notable being Grok 4.6 at $6/m which makes it a tougher sell

xmonkee 11 hours ago

It's really only between Anthropic and OpenAI for many of my use cases, since I have a Zero Data Retention agreement with both. I'm not trusting random inference providers and especially not Elmo with sensitive data.

speedstyle 8 hours ago

or Tinfoil [0]? They serve open models with container integrity attested by Nvidia/AMD enclaves. Every cloud provider offers this of course, but not usually in a way that can be shared between distrusting users for economical inference. It still relies on the open-source containers being secure, and there's probably hardware sidechannels and stuff, but personally (ie privacy not liability) I trust it more than a contract

[0] https://tinfoil.sh

Qwuke 6 hours ago

c0rruptbytes 9 hours ago

plenty of companies offer ZDR and are just as random as OpenAI and Anthropic in their age

xmonkee 8 hours ago

dataplumb3r 8 hours ago

You do not have a ZDR with Anthropic.

For Mythos and even Fable they require prompt retention on their end.

edit: or more precisely if you want to access Mythos/Fable ZDR does not apply, and depending on config the exclusion can affect other models.

OutOfHere 11 hours ago

Since when does Grok 4.6 have Sol 5.6's intelligence? I don't believe it.

maxdo 9 hours ago

I do have free sol and cursor ultra for 200 I prefer grok over sol, they are equally capable but grok is faster

qingcharles 5 hours ago

I use Grok 4.6 every day; it's good, but it's not Opus 5 or Sol 5.6. The gap is closing, though.

dimgl 10 hours ago

Why not?

chaos_emergent 10 hours ago

redox99 10 hours ago

It doesn't.

mohamedkoubaa 11 hours ago

I wonder if xAI is A/B testing routing some difficult grok 4.6 queries to Sol to seed some true believers.

bnrdr 9 minutes ago

dang, to avoid confusion from the title perhaps this should be edited to: “OpenRouter temporarily cutting GPT-5.6 Sol pricing by 50%”

dannyw 8 minutes ago

That suggests this is being funded by OpenRouter, and there's no indication this is the case (and I doubt OpenRouter can afford it; why would they anyway).

kelvinjps10 4 hours ago

I have switched to Chagpt sub now after only using Claude for coding. You get more value for your money and feels like codex has reached Claude code performance in coding (the reason for using Claude) regular plus account allows you to have access to their most powerful model, image generation and asking questions is better because you can use sol but in instant mode and it feels smarter and faster. And finally codex usage limits are better than the Claude daily 5h limit. And codex feels faster although Claude code had more features

andypants 38 minutes ago

> codex has reached Claude code performance in coding

Codex has always beaten claude in coding benchmarks, hasn't it?

gb2d_hn 4 hours ago

I switched as I felt Codex was on a par with Opus, but the chat responses from Sol are just more intelligible than the word soup I've been getting from Opus. I wonder if Opus could be prompted to respond in simpler prose via agents.md

FluffyPancake 3 hours ago

I have also noticed that I am increasingly struggling to read what LLMs are writing, finding it incomprehensible half the time. I stole Matt Pocock's line of "When reporting information to me, be extremely concise and sacrifice grammar for the sake of concision." for the agents.md It makes it a little better. I also specify to use https://github.com/AminBlg/SimpleEnglish/ for all writing it does including code comments. It all feels like a bandaids but that seems to be the best we can do right now.

marcyb5st 4 hours ago

Ah, so I am not the only one struggling with deciphering Opus writing style. At times I feel dumb as a rock because I read the same passage like 5 times and I still don't get it.

jen729w 3 hours ago

> I wonder if Opus could be prompted to respond in simpler prose via agents.md

God knows I've tried. I've got a variant of the ASD-STE100 trick which does the job, mostly, at the start … but get to about 100k of context and it goes out the window.

The model's personality is too strong for simple suggestion, alas.

CompoundEyes 10 hours ago

I used over a billion tokens per day of gpt-5.6 sol xhigh starting last Wednesday through Sunday before reaching my reset limit. The $200 pro plan is still the best deal.

ec109685 9 hours ago

I spent $800 in a few hours when my sub maxed out because I was trying to get something done and had a long car ride to let it churn.

Their api pricing is absurdly expensive.

giancarlostoro 9 hours ago

> Their api pricing is absurdly expensive.

I assume at this point that it subsidizes subscriptions.

timClicks 8 hours ago

HeWhoLurksLate 8 hours ago

qingcharles 5 hours ago

paxys 9 hours ago

Because API pricing is for corporations and subscriptions are for consumers.

wmf 6 hours ago

blitzar 3 hours ago

XCSme an hour ago

I am also on the pro $200 plan, the limit is high enough to do what I want for a week, and binge run Ultra Fast last day to use the remaining credits.

SJMG 10 hours ago

A billion a day? How many agents are you running?

_zoltan_ 3 hours ago

with ultracode, it goes fast. I can easily get to a billion on a busy day.

greenavocado 8 hours ago

That's wild. I was gonna say 3-5 billion a month is more reasonable summed across all token types.

infinite_spin 10 hours ago

I have mine churning like butter and I'm rarely hitting a billion tokens per day, what's your workflow look like?

CompoundEyes 9 hours ago

Pretty basic. The codex app with one conversation per project and several running simultaneously all hours. I’m going for max caching that way and it never gets lost even with compaction somehow. Each has a plan with milestones to keep up to date and a thin agents file. I check in on them in the Remote app. Use case is protocol and control reverse engineering of audio hardware. I think they must be identifying the heavy use agent sessions and cranking up their cache lives so it’s not a big deal for them.

childintime 36 minutes ago

rpdillon 9 hours ago

ramraj07 9 hours ago

Some people just do crazy stuff. For example this now ex yc guy who said he has agents constantly scanning Sf govt apis and forming dashboards just because

jimbob45 9 hours ago

dyauspitr 9 hours ago

I’m exclusively using ultra and I run out in 3-4 days consistently. Those resets are great but I’ve noticed they like to cluster them at the start of the cycle, would be better if they spaced them out more.

IshKebab 9 hours ago

A billion tokens per day?? Plausible estimates put the energy use at about 0.001 Wh/token, which means you're using 1000 kWh/day in electricity, just to generate slop. That's about the same as 50-100 houses. 300kg of CO2 per day - roughly the same as flying from London to New York every three days.

I think on average AI energy usage is not as big a deal as everyone is panicking about, but your usage is truly absurd and I don't know how you can live with that. It's immoral.

kolinko 2 hours ago

That’s for i/o tokens, mostly output. 90-98% is cache read usually, so you can divide electricity use by 10 at least.

As for co2, it depends on the provider, it could be way lower as well.

As for ethics, you don’t know what he works on, and how effectively - he might be saving 10x that much of co2 for the planet.

WASDx 3 hours ago

I'm glad someone is voicing this. Overconsumption at that level is not defensible. However if they meant cached tokens so it's not that bad.

Gangway0829 9 hours ago

I can't speak for that guy, but I'm a physicist and work in clean energy... So it's not too hard! That said, I usually am closer to 10M on days I do heavy coding, so not nearly that bad.

paxys 8 hours ago

Token caching is a thing

throwup238 8 hours ago

You really think OpenAI is selling $1500/mo of electricity (at $0.05/kwh) for $200/mo?

I’m guessing that Wh/token estimate is several orders of magnitude too high.

Noaidi 8 hours ago

lannisterstark 8 hours ago

1. You do not know what they're using it for. 2. Get off your high horse please. 3. Immoral my ass.

nozzlegear 5 hours ago

pertymcpert 5 hours ago

Most tokens are cached.

_zoltan_ 3 hours ago

"slop"? come on. we're not in 2020 anymore, Dorothy.

jm4 10 hours ago

I can’t sign up for that. I tried authorizing Codex a couple days ago. For some reason, their system says my phone number has been used for verification 3 times even though it definitely has not. I’ve had this phone number for over 20 years. OpenAI support is useless. They just keep repeating the policy without actually helping me.

weakened_malloc 9 hours ago

Use TextVerified, load up like $5 of credit and OAI verification is like $1.00. Then when your account is made, ensure 2FA/passkey is setup then you don't need to worry about the phone number.

stavros 3 hours ago

_345 8 hours ago

Yes I filed a support ticket with them and explained that their system is broken and they just did not care. I explained how it was impossible for me to use it 3 times already as I've only made 2 chatgpt accounts EVER, and only recalling entering my phone number for one of the two chatgpt accounts. I told them that this issue locked me out of codex and chatgpt for work and they weren't willing to do anything about it. Totally useless support.

I ended up borrowing my gf's phone number just so I could get access for work. Ridiculous

jaggederest 9 hours ago

Get a burner and use it? If you're spending $200/mo on something, $40 or whatever for a burner phone seems like a pretty cheap price.

63stack an hour ago

paxys 8 hours ago

ronfriedhaber 15 minutes ago

Hard to estimate what enabled the price cuts, Yet OpenAI is doing some magic work, especially recently.

cmiles8 an hour ago

This is the opening salvos of an all out token price war.

With models a commodity at this point there isn’t much leverage for the big labs to keep their pricing anywhere near where it’s at. And that’s at the worst possible time as they need to be dramatically raising prices to have a viable business model.

Expect pricing to rapidly fall towards the underlying cost of compute and as players get really desperate we’ll likely see inference at less than the cost of compute as the market starts to rationalize and squeeze out weaker players who’s only play left will be to be the cheapest option in town.

The AI bubble is just waiting for the first player to scream mercy and cut capex as they simply can’t afford to throw more cash on the burning pile. That will be the trigger that implodes this bubble.

z_rho_one 10 hours ago

If they can cut the price of Sol by 50% and the price of Luna by 80%, then the original price might have carried a massive operating margin. They might still be serving the models at a profit after these price cuts, but we will never know.

paxys 10 hours ago

I don’t think there’s a real answer for this. Margin depends on whatever number the accounting department wants to make up.

Do you include research and training costs? Of all models or only the ones being served? What percent of the R&D budget do you allocate to inference? What about data center capacity? Do you count future commitments? All the circular financing deals? Do you count employee equity grants as costs? At what valuation?

dannyw 3 minutes ago

We have a simple definition for this: COGS.

We also have another solution for "whatever accounting decides": generally accepted accounting practices. It's far from perfect, but GAAP figures are what you should be looking at; not "adjusted GAAP" or whatever invention.

kolinko 2 hours ago

Or they have gotten new asics and can do now inference way cheaper

XCSme an hour ago

Why would they infer faster with new shoes?

wahnfrieden 10 hours ago

OpenAI didn't cut the price of Sol by 50% like they did with Luna's 80%. Sol was unchanged. This is just a limited promo for OpenRouter non-BYOK.

InsideOutSanta 4 hours ago

I'm pretty sure tokens are priced to maximize revenue, not inference profit.

josu 5 hours ago

I always find it funny that Japanese pensioners are probably subsidizing my tokens.

dgellow 2 hours ago

What do you mean?

josu an hour ago

stillpointlab 4 hours ago

I like to see this. I still prefer Fable (marginally) but my last big task was 100% Codex using Sol max (re-sizing my AWS infrastructure using CDK) and it did a very good job. No complaints, I could use this model happily to do what I need to get done.

If this nudges Anthropic to give me more Fable usage, that's even better.

0xbadcafebee 4 hours ago

> my last big task was 100% Codex using Sol max (re-sizing my AWS infrastructure using CDK)

Fwiw, you could do this with any small or medium model, and it's easier with the aws-docs mcp. AWS is pretty stable, well documented, and programmatic, so most AI can figure out what it needs pretty quick

claiir 4 hours ago

Since it's only discounted on the standard "OpenAI," non-ZDR route (old pricing on Azure), I'm guessing a lot of users won't see this benefit? Since a lot of users enable a global "ZDR-only" toggle on OR

stavros 2 hours ago

It would seem that getting lots of data is exactly the reason to discount this.

dannyw a few seconds ago

OpenAI says they don't use any API data for training.

m4rtink 10 hours ago

Price wars did wonders for many businesses, like the bike sharing industry in China.

Overgrown datacenters or mounds of GPUs dumped into the harbour next ?

Moto7451 10 hours ago

I would in such a scenario expect the GPUs to be dumped to industrial breakers who would send them to China for refurbishment and repackaging before being sold again on Amazon, AliExpress, and Taobao as last gen gaming cards from weird brands and specs.

This is what happened after the great crypto GPU dumping.

dawnerd 7 hours ago

Honestly can't wait for that to happen, same with memory, drives, etc. There's going to be a massive amount of server pulls hitting the market.

Joel_Mckay 5 hours ago

The e-waste recyclers are pretty low on the pecking order, as the creditors will be first to strip these places for assets as Leopold Aschenbrenner discovered. =3

m4rtink 10 hours ago

Yeah, I ment it as a joke - I agree with you. Watched the Gamers Nexus GPU investigation recently, where they were shown how a chinese soldering shop can transplant GPU chips to a new board, including memory chip reuse.

Hopefully we can look forward to all that useless datacenter AI crap gets repurposed in a similar manner into something actually useful for users.

kajaktum 10 hours ago

skohan 5 hours ago

I would happily buy up a load of datacenter GPU's at deep discount

Joel_Mckay 5 hours ago

Won't have to wait very long... as they are eating their own already.

https://www.youtube.com/watch?v=rE75WvOtcu8

The Shrek movie market correction correlation may be due again in July 2027. =3

throwatdem12311 7 hours ago

Mountains of GPUs next to the ET games in the landfill.

vatsachak 8 hours ago

Well one person can use at most one bicycle at a time.

One person can use as many GPUs as they want.

krzyk 5 hours ago

Is this pricing change only for openrouter? I don't see official OpenAI info about this.

Bombthecat 4 hours ago

I was wondering the same, and it clearly says : 50% off, aka a sale, not normal price cut.

I don't get this thread.... Really. Is it full of bots?

egorfine 2 hours ago

Slightly unrelated: what's up with the "tps" value? Does GPT-5.6 Sol really deliver just 32 tokens/second?

ardel95 2 hours ago

My bet is that OpenRouter began steering GPT-5.6-sol users towards flex tier, which is already 50% off.

So this isn’t really a price cut. As to why, lots of possible reasons. Perhaps an agreement with OpenAI to help them drive up more diverse traffic priorities.

matheusmoreira 5 hours ago

Does this mean less subscription credit usage as well?

kaycey2022 4 hours ago

No because i am still losing 50% of my weekly quota using sol on high.

Tadpole9181 5 hours ago

It just looks like OpenRouter is doing a 50% sale on some models right now? Until OpenAI makes an announcement, I would assume no.

0xbadcafebee 4 hours ago

OpenRouter doesn't do sales, they charge a premium, which is a flat 5.5% taken out of your credits. If you see something cheap on OpenRouter, it's because that one provider lowered its price. (Actually, correction, they will take 0.5% off their fee if you allow them to train your content)

Here are all the providers giving discounts: https://openrouter.ai/collections/discounted-models

Another thing some people don't notice is flex pricing, which is way lower than default pricing, for slightly worse latency and reliability. Depends on the provider and model

throwatdem12311 7 hours ago

At this point the models are “good enough” and whoever wins long term is gonna be whoever is the cheapest.

That’s why Chinese models are gaining traction and it’ll be the only way for OpenAI or Anthropic to keep up.

dgunay 10 hours ago

I'm loving this race to the bottom.

infinite_spin 10 hours ago

I'm not having that experience. So far each major model update has been at least slightly better than the last, in ways I've found useful. Can't say it's perfect, or able to do exactly what I want without a decent amount of instruction/implementation/docs, but it's been useful enough to keep paying for it.

fn-mote 9 hours ago

GP means race to the bottom in price not quality.

dgunay 7 hours ago

Oh no the models are absolutely getting better, I'm just amazed that only 6 months ago I was using gpt-5.3-codex, and now I can use gpt-5.6-luna for similar results at like 1/15th the cost. Now 5.6-sol is being slashed by 50%? Amazing.

Taikhoom10 8 hours ago

josh-wrale 11 hours ago

Is this motivated by the value of the thinking traces gleaned from the traffic?

killingtime74 9 hours ago

The thinking traces are server-side, not exposed

ec109685 9 hours ago

They can’t decrypt the thinking traces.

dannyw 8 hours ago

You can train a LLM to inverse summarised thinking into thinking text. It’s not perfect, but it gets you maybe 80% of the quality with proper techniques.

Paper: https://arxiv.org/abs/2603.07267

FWIW, there’s not that much value protected here anyway IMHO, and even raw thinking text can lie (as shown by Anthropic’s amazing research), so for legitimate interpretability research it’s limited.

Scaling frontier performance hasn’t been SFT-bounded for a while now; it’s now basically how much you can scale RL rollouts.

OutOfHere 11 hours ago

The title looks to be misleading, since this price cut is limited to OpenRouter. It does not apply for the native OpenAI price listed at https://developers.openai.com/api/docs/models/gpt-5.6-sol

jrflo 8 hours ago

Yes, and it's only for a month. This is an ad.

paxys 10 hours ago

Which raises the question - who is subsidizing this, and why?

internetter 10 hours ago

Possibly OAI? If you have OAI tokens you are a captive audience. If you have OpenRouter you are bidding on a free market.

OpenRouter attributes this promotion to OpenAI https://x.com/OpenRouter/status/2089416739398254662

maxnevermind 5 hours ago

paxys 9 hours ago

OutOfHere 9 hours ago

OpenRouter is likely just leveraging Codex subscriptions.

prime_ursid 9 hours ago

matchagaucho 10 hours ago

Right? Should we switch from direct OpenAI API integration to OpenRouter?

What's the incentive here?

Open Responses API doesn't appear to support state management (yet)

user43928 4 hours ago

OpenRouter and the Vercel AI Gateway.

So yes, presumably a very small share of their total traffic.

lyjackal 9 hours ago

I saw this for Luna and then looked at the uptime and it said 85%. My interpretation is that this is just a gimmick where they serve the OpenAI flex tier at the same discount OpenAI provides for flex and then fall back to azure

drivebyhooting 8 hours ago

Has anyone had mixed experience running Ultra with and without /goal? I come back to it after 8 hours to find it got stuck navel gazing imagined and Byzantine errors.

tartakovsky 10 hours ago

No ZDR. No dice.

hk__2 2 hours ago

In my experience, "Sol" stands for "Stupid overengineering LLM". I’ve tried it at low/medium/high/xhigh effort levels and after a while I always end up to regretting my switch from Opus/Fable.

therepanic 8 hours ago

Even at these prices, switching from subsidized subscriptions to the API just isn't worth it. Not even close.

bigbluedots 3 hours ago

These threads seem to have become exceedingly vibes-based. Yes, something may now be cheaper or more expensive or whatever, but there is no way to objectively measure quality (except for "trust me bro" benchmarks). So the discourse is people saying that for them, this or that model was better - which is a very low value data point.

oblio 3 hours ago

> These threads seem to have become exceedingly vibes-based.

If you remember programming language discussions, they are exactly like this.

Software development is still in the leeches and bloodlettings phase.

bigbluedots 3 hours ago

Yes, programming language discussions can be vibes-based and therefore low value too, but sometimes the more concrete aspects of the languages at hand, e.g. language features and tradeoffs are discussed. That is something that I'm not seeing in equivalent AI discussions.. there is a lot of how a particular AI model "feels" to interact with.

jeffybefffy519 7 hours ago

Reading the comments in this thread, i honestly dont get it. 5.6-sol has felt like a regression in capability. In fact, every model since 5.3-codex has been a regression from OpenAI. I just find 5.6-Sol over engineers problems, takes absolutely ages to solve basic problems....

At this point, I'm considering going back to cursor over codex due to the ability to get more control over what model I use since there is clearly a heap of user preference and having frontier providers constantly shift the goal post with "State of the Art" is complete non-sense.

jeswin 7 hours ago

It depends on what effort you're using etc. As an example [1] of what codex is capable of, here's hugo (written in golang) ported to TypeScript - and then a TypeScript to Rust transpiler which converts arbitrary TypeScript into Rust.

The TypeScript code which was transpiled into Rust (and is compatible with most hugo templates) runs faster than the original hugo.

[1]: https://github.com/tsoniclang/tsonic-examples/tree/main/rust...

The transpiler is still WIP, but the fact that it can do this says a lot of about how far LLMs have come.

jeffybefffy519 39 minutes ago

I find all effort levels of sol are the same in terms of amount of hallucinated unnecessary changes. Luna is much better all round on xhigh but my point still stands, every release of these new models is not an upgrade, its re-learning how to work with it.

Its like rehiring an employee every few months then training them up. Its honestly tiring and cant stay like this.

Opus has the same problem too…

jeswin 30 minutes ago

tonyhart7 7 hours ago

its over engineered problem solver ???? well because its a designed to do that

if you want to solve basic problem then use Luna

jeffybefffy519 41 minutes ago

I mean it added additional changes when it doesnt need to. Its basically hallucinating changes it thinks it needs to make regardless of effort levels i try.

shevy-java 5 hours ago

They are really getting desperate. The bubble is coming closer to an end here.

ComputerGuru 10 hours ago

Does OpenRouter eat this cost to get their hands on a copy of the conversations people are using with the model?

9cb14c1ec0 9 hours ago

No, this is OpenAI doing the discount, not Openrouter by themselves. OpenAI is crushing it with their 5.6 models, and they probably decided there was no better time to grab as much market share as possible.

Noaidi 8 hours ago

I don’t understand this at all. They have never been profitable yet. How is this helping them? When it be more likely the case that not enough, people are using it as the prices they established already? So now they have to lower the prices?

ipaddr 6 hours ago

nprateem 6 hours ago

voiper1 6 hours ago

OpenRouter offers 1% discount to save your conversations, explicitly opt-in.

Anything else they don't save it. Even if they tell you the model provider saves your data for training.

dvrp 9 hours ago

For context, Stripe has just acquired OpenRouter for >$7B.

I’d bet that explains this move!

indigodaddy 9 hours ago

Why would the potential acquisition have anything to do with this? They do discounts all the time on various models. Luna was 50% off last week..

gip 9 hours ago

Not sure as OpenAI models (Sol, Luna,..) are also discounted on the Vercel AI Gateway rn. My bet is on OpenAI trying to drive more enterprise customers to their models through API.

Topology1 7 hours ago

How can they do this? Are they subsidizing it out of pocket?

ben8bit 5 hours ago

Terra is also a fantastic model.

gutterscale 4 hours ago

GPT-5.6 sol starting to be a real workhorse at this price point

kristo 2 hours ago

It shocks me how little people seem to care that they are supporting an evil Zionist lizard man who molested his sister and is happy supporting trump. Doesn’t even come up in the conversation here. I don’t really care if sol is a bit better, I still make decisions on more than that.

Is the HN community just too online and sucked in to the musk mind manipulation vortex? Or what is going on? Why does nobody seem to care?

aetherspawn 5 hours ago

Can we get it for the reduced rate direct from OpenAI though?

vorpalhex 11 hours ago

Do other people find 5.6 to be worse at most simple tasks and frequently over complicate things?

I asked it to write a user todo and it turned out a four page essay. I gave the same task to 5.4 and got the small list of checkboxes I expected.

qup 10 hours ago

I've found it to be great for planning code changes (or new projects). I use the superpowers plug-in which I think guides the planning.

Then I switch models (to luna) before implementation. I find this combo nearly always does what I want.

I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above.

I like that skills are easy to try out, too.

jeremyjh 7 hours ago

I stopped using superpowers because it wanted to turn every tiny bug fix into a $37MM DOD project. I got effective results but it took ages. I may try again - I need to find a good way to run different profiles in my harness so I can easily shut it off. The default planning workflow in OMP is pretty good though.

I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if the task is well defined and straightforward, but there are lots of models now that you can say that about.

jcastro 8 hours ago

I have the same setup you have, love it!

drdexebtjl 10 hours ago

You would probably get better results with Luna for the real simple tasks, or Sol with low thinking effort.

I find that I get exactly the effort that I asked for, which is pretty nice. The other side of that coin is that these are the least lazy models I’ve used so far. They will go on elaborate tangents to complete the task when I want them to.

khacvy 7 hours ago

any good best practice for effort selection on claude. I always use high as default.

infinite_spin 10 hours ago

I've found it's worse for simple tasks too, and I have to give it stricter guidelines, and sometimes it doesn't follow the same patterns I've grown to expect. I've found using 5.6 (sol) is good for diagnosing issues though, especially in terms of optimization of some given path

dimgl 10 hours ago

Yep. I have not yet had a single good experience with Sol or the 5.6 models on a variety of harnesses and configurations. It overthinks, overcomplicates and often makes my code into an unmaintainable sludge. It'll usually take 5+ turns of steering to get it in the right direction.

OutOfHere 11 hours ago

It's your responsibility to set an appropriate level of Thinking. For simple tasks, I use the instant model. As an approximation, the choice is proportional to the amount of time I want it spending on the task. Also, you can always ask it to respond succinctly.

gxs 7 hours ago

Absolutely not

I’ve used Claude exclusively for the past few months

Was excited when Sol came out a few weeks ago and loaded it up

I made the mistake of treating it as if it were Claude - I’d assumed they were close enough in ability and treated them that way

Well, turns out my instruction sets for Claude are 100% too complicated for Sol

Sol made the stupidest assumptions, constantly did things that it wasn’t asked to do and always approached code in what I considered a weird way - I had redo a lot of my prompts to get it anywhere close

Now, did it do good work?

Yes, on occasion. But with LLMs and coding, consistency is the name of the game. Constantly having to correct the LLM and constantly feeling paranoid that it won’t listen makes for an exhausting session

Maybe if you “came up” in the codex world you’re more fluent with it, but sticking with Claude for now

dana321 3 hours ago

"rewrite unreal engine in rust, make no mistakes"

iammrpayments 6 hours ago

Why are you being downvoted, is this post an ad or something.

cbg0 5 hours ago

Probably because it's PEBCAK.

gxs 4 hours ago

Scene_Cast2 8 hours ago

Oh hey, that's cheaper than Kimi K3! Amusing to see a SOTA OpenAI model be cheaper than a Chinese open weight model.

Fwiw I love K3 and use it as a daily driver. I haven't tried Sol, as I dislike OpenAI.