Grok 4.7 (x.ai)

434 points by meetpateltech 6 hours ago

moojacob 6 hours ago

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

imron 37 minutes ago

> My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish.

Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.

It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.

I find I often have to ask it to re-explain what it means.

iamflimflam1 9 minutes ago

It’s frustrating that we can’t see the “thinking” - it’s like we only have access to half the conversation.

smashers1114 5 hours ago

FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.

a2dam 3 hours ago

I think this is more a meme than anything else, for a couple reasons:

First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.

I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.

Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.

smashers1114 3 hours ago

TuxMark5 3 hours ago

oxidant 2 hours ago

StilesCrisis 3 hours ago

_boffin_ 4 hours ago

Does not work for Claude, at least for me and I put it as the system prompt

bel8 3 hours ago

tempest_ 2 hours ago

LPisGood 4 hours ago

ffsm8 4 hours ago

junon 21 minutes ago

I'm wanted to try this exact thing! I'll have to try this now.

BatteryMountain 2 hours ago

I tell mine to address me as a tech priest of the adeptus mechanicus. Works great.

SoMomentary 3 hours ago

I created a custom output style based on this (borrowing some from github.com/AminBlg/SimpleEnglish) and I've found it to be better than the default or concise output styles, but still not as good for me as current GPT or Gemini models when it comes to communicating.

el_benhameen 3 hours ago

I tried this a while back and I felt like the result was the same weird shoehorning of ideas into language, just with a different vocabulary. I’d really like for it to work, though.

snapplebobapple 4 hours ago

This fixed claude! Thanks!

guluarte 2 hours ago

I put this rule in my CLAUDE.md: "Always write a TLDR in layman terms", it seems to do the trick

jasonjmcghee 6 hours ago

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

vessenes 5 hours ago

I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.

boc 2 hours ago

I've been getting a ton done with Fable as the supervisor and astra as the implementer, with opus for adversarial reviews of the astra PRs. You can use terminal multiplexers with custom harnesses to allow Fable to start codex sessions and send instructions / read instructions / allow/deny actions. It's pretty cool!

jitl 21 minutes ago

vintermann 5 hours ago

It's not just about benchmaxxing. Sincerely targeting those long-autonomy benchmarks is questionable in the first place, because naturally it drives the model to assume more and more about what you want.

svachalek 4 hours ago

Lucasoato 5 hours ago

> I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

TomGarden 5 hours ago

Agreed. The more knowledge you amass on a subject, the more important it becomes to be extremely specific and nuanced - or your communications end up being incorrect. You become better at expressing your thoughts, but harder to understand.

The weird thing is, that's not what AI models seem to be doing. The prose is just weird.

unshavedyak 5 hours ago

fearmerchant 3 hours ago

MisterMunchkin 3 hours ago

samuelknight 5 hours ago

That's half true. A very smart model should be able make good explanations, which include simple understandable prose. That can should be possible even as its thought process gets more alien.

Aperocky 5 hours ago

The best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.

When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.

fragmede 5 hours ago

superjan 5 hours ago

What I notice about Claudish is that it has its preferred cliche’s and overstretched methaphores, it packs too many ideas in a sentence, and to achieve the latter it makes up adjectives.

I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.

svachalek 4 hours ago

yread 4 hours ago

Yeah just today it told me in a snarky way that my CPU (7940HX) doesn't exist and that I must have misread it and it's either 7945HX or 7940HS. Yes, AMD (re-)branding CPU models makes things difficult but I thought we are past AI models making such egregious mistakes

thesmtsolver2 4 hours ago

This is /r/iamverysmart material (by Claude)

Part of intelligence is knowing your audience and communicating efficiently.

cruffle_duffle 4 hours ago

grababner 5 hours ago

If you can't explain it simply, you don't understand it well enough

WarmWash 5 hours ago

Perhaps you haven't had the chance to use it, but 3.8 flash is the best model for talking too. Even routing Claudes output through 3.8 to have it explain whats going on is a breath of fresh air

svachalek 4 hours ago

Agreed. It's very capable for something carrying the "flash" label, super fast, and very clear to read.

moojacob 4 hours ago

I'll have to try Gemini Flash for coding. The reason I haven't I used Gemini for coding is last time I tried it couldn't call tools very well.

I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.

haellsigh 3 hours ago

WarmWash 3 hours ago

AustinDev 5 hours ago

gemini 3.8 flash?

jtwaleson 5 hours ago

esafak 4 hours ago

I would if they let me bring the subscription I have to the harness of my choice.

johnsimer 4 hours ago

I've found grok 4.6 speaks heavily in Claudish. It especially likes using verbs as nouns.

giancarlostoro 2 hours ago

> Claudish

I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.

dumberquestions 5 hours ago

Token price doesn't tell you much without knowing token efficiency.

user43928 5 hours ago

Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

How representative that is of real world usage, I don't know.

In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

attentive 3 hours ago

$0.50 for cache reads, which is 25% of input. While other models are 10% of input.

And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).

tk90 4 hours ago

> I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.

dmix 4 hours ago

Agreed, Claude has a "Claude Design" tool but doesn't publish any frontend brenchmarks. Maybe the industry will develop one.

rayiner 4 hours ago

> My favorite part of the new Groks has been how they speak in plain english.

I don't know if it's the plain english or what, but I really like Grok for legal research (as opposed to code). It's got a noticeable edge in getting to the point compared to Opus 5.

qaq 2 hours ago

For me Grok finds legit bug that Fable and Astra miss so I always run it as part of code review

algoth1 4 hours ago

I've noticed Chatgpt 5.6 Sol High, on the chat interface, inventing words that are a mixture of Portuguese and English. Like "hardcodar" a mix of "hardcode" and the most common verb ending in Portuguese "-ar". Some don't have a single google hit

shawabawa3 3 hours ago

Do you have any connection to Portugal? I imagine if you have Portuguese in any of your prompts that might bleed into your user profile which becomes a part of every prompt. Alternatively it might use browser language settings

Waterluvian 5 hours ago

Using a variety of models feels similar to the benefit of having a team of individuals from different backgrounds.

pietz 4 hours ago

Looking at AA and Vals, your theory seems to check out.

atniomn 5 hours ago

I expect the next Anthropic release to finally reduce the prevalence of Claudish

imron 43 minutes ago

I expect the reduced prevalence of Claudish will have its own mannerisms that become the new Claudish.

The Claudish is dead. Long live the Claudish.

moojacob 5 hours ago

If they fix Claudish, they've earned me back as a max customer!

Fable 5.1 is not there quite there yet.

They need to get that Sonnet 3.5 magic back.

rfgplk 5 hours ago

sscaryterry 5 hours ago

Based on?

7734128 5 hours ago

fatata123 5 hours ago

xmorse 5 hours ago

it's definitely not bigger. smaller if anything looking at how much faster it is

petesergeant 3 hours ago

Grok and Zai have both been excellent as adjunct code-reviews, on their cheapest plans, for me. Fable plans, Opus writes, Codex as primary reviewer, but Grok and Zai usually find something worth fixing that the others have missed. Both are well worth whatever the $20 or so I'm paying for them

Forgeties79 3 hours ago

I do not understand how anyone can seriously use a tool that has "Be funny and irreverent when appropriate" baked into the system prompt.

I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.

StilesCrisis 2 hours ago

I am on the exact same page as you, but there is definitely a market for LLMs which speak more conversationally and less like Claude! Non-programming use cases abound and most users don't like the rigid, exact tone that engineering demands.

mchusma 18 minutes ago

Initial impressions, Grok 4.6 for me just didn't really hack it for any usecase I tried. I seem to have a floor for my usecaseses (coding and a bunch of agentic workflows) and Sol/Opus are above some kind of intelligence floor.

4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.

Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.

andsoitis 17 minutes ago

> for any usecase I tried

For example?

mchusma 10 minutes ago

Coding, agentic flows like logging into my accounts and gathering data, grok bot.

4.6 made more mistakes than SOL or Opus overall. Gave up a lot. And in my opinion, the rate of mistakes is kind of more important than how brilliant it is.

I think 4.7 may still be better, but I was hoping for clearly Sol/Opus level and so far it just isn't there for me.

vessenes 5 hours ago

Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!

jjcm 2 hours ago

It's definitely gotten better at image->html workflows. Here's a test comparing Astra (currently SOTA at this) vs Grok 4.7:

Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

Astra's build: https://html.non.io/annui/

Grok's build: https://html.non.io/Annui-grok/

Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers."

Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit.

Still though, for the price it's a great starting point.

jjcm 2 hours ago

For transparency, it took around 20M cache read + 1M input + 100k output tokens. API pricing for each:

Grok 4.7: $12.60

GPT Astra: $35.00

simonw 5 hours ago

https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.

Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.

UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.

For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

nicolamanzini an hour ago

Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.com

Grok 4.7 generations: https://threejseval.com/models/grok-4.7-high

Also go vote on https://threejseval.com so you can help evaluate how Grok and other model performs compared to each other!

TomGarden 3 hours ago

I think these are the worst I've seen, at least in some time. It's a silly benchmark though, not sure what to make of it

athrowaway3z 2 hours ago

I think the result is fine. The benchmark is silly to the point of being useless nowadays.

It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.

So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":

Does the user want the least lines of code to make it functional, or the best looking version?

Mashimo 2 hours ago

If you think this is bad, look up mistral.

forgot-my-pw 4 hours ago

I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=co...

datsci_est_2015 4 hours ago

Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.

MattDamonSpace 4 hours ago

Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file

kiliancs 4 hours ago

What is the default reasoning level?

daveguy 2 hours ago

Hahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?

simonw 2 hours ago

saejox 5 hours ago

Not even close to astra. Astra is something else. It is expensive, but uses way fewer tokens do my tasks.

xAI missed its chance, Ball is on Anthropic's court.

redox99 4 hours ago

Not surprising considering Grok 4.7 is a 2T model, so Sol/Opus class, not Astra/Fable class.

jstummbillig 3 hours ago

How many parameters do Astra or Fable have?

brianwawok 2 hours ago

01100011 4 hours ago

I tried Astra w/ high reasoning on a design document project and it was horrible. It started duplicating output lines, made document edits without permission, and basically did a poor job writing clear prose. I went back to 5.6-sol and it's great. I'm an OpenAI fanboy and was severely disappointed. I hope Astra is better for coding.

manmal 3 hours ago

No, Astra isn’t better for coding. I’ve switched back to Sol.

Razengan 2 hours ago

enraged_camel 4 hours ago

Astra fails in similar ways, and at similar frequency, as GPT 5.6 Sol does. It often goes way out of scope, or just stops prematurely, or tries to find odd and even dangerous workarounds when it gets stuck.

It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.

sneezychl 28 minutes ago

LLM's introduces problems, and it finds them in its own internal thinking. But instead of actually modifying the previous generated answer to fix the real issue, it adds another layer to deterministically guard around it, greatly expanding the scope of the fix. This scales with effort, and the result is spaghetti and with a side of bugs.

Best to stick with a high end model + low effort, do a manual pass on high effort and fix the bugs you know are reachable.

brink 4 hours ago

Same, Astra is extremely RL fried, and nobody is talking about it. I used Astra for a few days on my personal project, and load times went from less than 3 seconds to almost 30 seconds because it kept using the wrong sync primitives and bad architecture overall.

haellsigh 3 hours ago

epolanski 2 hours ago

I don't understand these comments.

The two models are in completely different price tiers. Astra costs 5 times as much.

It seems like all you can judge about cars would be their maximum speed on an oval.

user43928 2 hours ago

If you have a look at their headline benchmark on the post here, Grok 4.7 is hardly cheaper than Fable 5.1 Low and performs similarly.

Based on Artificial Analysis Cost per Task, Astra is about 2-3x cheaper than Fable 5.1 at Medium and Low.

Consequently Astra could be cheaper than Grok 4.7, depending on the task.

qwerpy 4 hours ago

I've been using 4.6 for some one-off game mods/utilities and it has done very well. "I have a very niche keyboard (Moonlander) and I play this very niche space sim, make me a SVG keyboard cheatsheet for it". Told me to grab keymap.c for the keyboard and inputmap.xml for the game's key bindings, churned for a while, then spit out a pretty good first attempt, along with the python script used to generate it. Spent another hour of back and forth to refine the script, and now it generates great diagrams that will adapt as my keyboard firmware and game bindings evolve: https://files.catbox.moe/x0u76x.svg

Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.

Theodores 3 hours ago

Impressive! I had to peek at the SVG file and it superficially looks good, however, as is the case with everything AI, the more you look, the more it doesn't make any sense.

By now AI should know of the DRY concept. But no. Hence the keys have a rounded rectangle for the key shape and another rounded rectangle for a clip path, to prevent text overflow. There are 72 * 2 = 144 identical rectangles, when just one would suffice (in the defs), with this being cloned once for the clip path, and 72 times for the keys.

I would not expect SVGO levels of optimisation (rounding numbers, that sort of thing), however, the human, if writing out the same thing for the 72nd time, might think 'is there a better way', to get the manual out. A graphics program such as Illustrator would not do that, but AI 'should' because AI.

The above is not criticism of your work, just an observation regarding AI SVG capabilities.

meerita 5 hours ago

Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.

vorticalbox 3 hours ago

Compared to the deep seek, gml sure but compared to OpenAI and Anthropic it’s actually very cheap.

In cursor I have switch over to grok for planning a composer for coding.

brianwawok 2 hours ago

Maybe mid priced is a better term for it lol.

vorticalbox 2 hours ago

thefourthchime 3 hours ago

It’s a great value if you get Cursor Ultra. I basically have infinite tokens

testfrequency 5 hours ago

What is the most secure way to use this model as someone who is lazy

user43928 5 hours ago

I understand DeepSeek 4.1 Flash is available on US providers with Zero Data Retention if that is what you are asking.

brcmthrowaway 4 hours ago

sparkling 4 hours ago

drewnick 3 hours ago

I use it on fireworks which is US/ZDR and pretty reliable. We run a few hundred million tokens/day through it for dollars. Many are cached, which is super duper cheap.

simlevesque 4 hours ago

I like devcontainers

parineum 5 hours ago

Brought to you by...

meerita 5 hours ago

By no one. For the price of 1M token you can get more and with better results with other models.

includenotfound 3 hours ago

_s_a_m_ 4 hours ago

DeepSeek 4.1 Flash is garbage, it almost only produced trash code. if you do extremely dumb things it is maybe sometimes fine to use.

yipinwong 4 hours ago

Not only that all DeepSeek is all garbage.

GLM or Kimi are better for my own personal projects. DS? uhm. it just keeps doing dumb crap

trentor 4 hours ago

Looks like they have still problems with caching. Prize is double the other providers for cache hits... which is most of what I do. :/

sarjann an hour ago

Why are they comparing grok 4.7 xhigh to grok 4.6 high?

Unless they produce the same token output on the face of it, it looks like they're trying to cover for 4.7 not having good model perf?

maz1b 5 hours ago

Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.

avazhi 5 hours ago

Your comment is like 6 months to a year late.

There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.

johnfahey 4 hours ago

No doubt xAI has seen rapid progress, but it's been several months of them being "just behind" OpenAI and Anthropic. It seems the gap between just behind the frontier and pushing it is a lot wider than most people thought it was a year ago, and that's why a clear third contender in the frontier model space has yet to materialize.

WarmWash 5 hours ago

Good thing they used 5.6 sol instead of Astra for benchmarks, the EEbench one is crazy[1]

[1]https://eebench.org/

notduckrabbit 4 hours ago

Significant regression in token efficiency compared to Grok 4.6 suggested by artificialanalysis.ai Intelligence Index Comparisons.

sourcecodeplz 3 hours ago

Output tokens from Intelligence Index:

- grok 4.6 (xhigh): 97M (for 44 score)

- grok 4.7 (xhigh): 240M (for 46 score)

everfrustrated 3 hours ago

That is comparing Grok 4.6 high to Grok 4.7 xhigh tho.

notduckrabbit 3 hours ago

No, you can add Grox 4.7 high to the chart. 36k vs 66k

GodelNumbering 4 hours ago

Every Grok release obscures their cache pricing while highlighting their input/output pricing

From their headline comparison:

  Grok: $2/$6 per million
  
  Fable: $10/$50 per million


  What this doesn't say: Grok costs 0.50/M cache read, Fable $0.25/M cache read
Long running agentic workflows are dominated by cache reads.

Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)

sourcecodeplz 3 hours ago

muse spark 1.3 contribs cache read is $0.002 btw (~220x diff).

bastawhiz 23 minutes ago

Is anyone treating Meta's offerings as a serious contender in any real use case? Zuck and co are burning cash hard to try to get people using their models after falling off the wagon for a couple years. It would be wild if those prices aren't total loss leaders.

c0rruptbytes 4 hours ago

as someone who is limited by amazon bedrock support at work (no idea why we got stuck with the worst one) - grok is literally the only budget-ish model option, so nice to see it updated, Sol and Opus are just too rich for my blood. Luna is good but so slow at getting things done (tps wise it's fast)

mh- 2 hours ago

Are the prices on Bedrock substantially different to the rate cards of the direct APIs? Just trying to understand whether this is Opus-through-Bedrock is too expensive, or Opus is too expensive.

gslepak 4 hours ago

Does anyone have any experience with Grok's subscription? How does it compare price-wise to the API?

daquisu 3 hours ago

There are some users reporting it improved a lot in the last few weeks. The max sub usage for Grok is around $12,000 of API pricing now, so a 40x multiplier for the $300 plan.

It is the same multiplier for Sol with subscription. For Astra though the multiplier is ≈20x, so half of Sol usage.

For Claude it seems to be ≈40x too for Opus, but less for Fable (similar to Astra in GPT).

All on the most expensive plan. Previously, Grok usage escalated linearly from the $100 plan to $300 plan. That would be a really good $100 plan if it is still true.

Some sources:

1. https://x.com/kunchenguid/status/2098256018836963382

2. https://x.com/stevenzhang/status/2092110386569089311

3. https://github.com/openai/codex/issues/43731

4. https://redd.it/1wciwc1

5. https://x.com/SemiAnalysis_/status/2064815044085318040

6. https://redd.it/1vx0k69

everfrustrated 3 hours ago

I find I can just about get by with coding every day on a Cursor $60/mth sub with Grok fast mode disabled. Doing pretty heavy coding work/requirements etc, but not much sub agents and no loops.

For me and what I’m doing that’s insanely good value.

I find grok build chews through my SuperGrok sub very quick - but I think that is due to it having the 500k context window which uses more credits. Cursor limits it to 256K (tho I see in today’s update for Grok 4.7 there’s now a toggle for context size).

thefourthchime 3 hours ago

There are two ways to subscribe, and it’s very confusing, but the best value is to get cursor ultra for $200 a month. I basically have infinite tokens with that plan, plus grok bot, which I really like

andreyvit 4 hours ago

Well when I ran out of Grok SuperHeavy subscription ($300) once and tried to use extra credits to cover half a day remaining till reset, $50 in extra credits went in two hours. Based on that, subscription definitely lasts longer; Grok subscription just about covers a week of my work (sometimes a bit extra remains unused, sometimes it runs out half a day to a day early). And as a point of comparison, it lasts for doing same tasks as 2.5-3 weekly limits of Codex on 5.6 Sol did (using xhigh on both Sol and Grok); I needed 3x$200 Codex subscriptions to cover my weekly usage.

nwienert 4 hours ago

By far the worst value subscription of any. I tried Superheavy and got about 5-10% the usage of CC/Codex.

ls1911 6 hours ago

after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai

shdtabasum 4 hours ago

Why Chinese models from Kimi, Deepseek are not added in comparison benchmarks?

xquce 4 hours ago

Same reason Coca-Cola only mention Pepsi and Pepsi only mention Coca-Cola. It's an proven way to capture the market. You would rather split the pie in two rather than in 4,12 or 50 right?

swalsh 4 hours ago

Codex has become my goto tooling. I used to be a Claude Max subscriber, but I was becoming disappointed with the quality of the output from Opus 5. Fable chewed through my usage too quickly to be practical. Moving to a Pro account w/ Codex was a big improvement. Sol had great output, and the usage was more than sufficient for most of my needs. However astra does tend to chew up usage, so when i've done to much of that, and it's became an issue Grok Build has beocme my second go to account. The output especially after the cursor purhcase has become quite good, and the usage has always been very generous.

becquerel 4 hours ago

Try using astra as an orchestrator for deepseek 4.1 flash, it seems to work out quite well.

sparkling 4 hours ago

I am using exactly the same flow.

Astra for deep dive investigations, Sol 5.6 at mid-level for day to day tasks, Grok 4.6 via Cursor for routine and low complexity tasks.

alansaber 3 hours ago

As anthropic/openai subscription allocations get squeezed you'll see more people using "second rate" closed models like grok. The token allowance with a Cursor subscription is crazy.

6thbit 5 hours ago

( why is the x-axis on the first chart in descending order ? )

simonw 5 hours ago

$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?

sejje 5 hours ago

cached input tokens are $0.50 per 1M (prompts under 200k tokens) and $1.00 per 1M (200k+)

simonw 4 hours ago

Do other prices vary for >200,000 or just the cached tokens?

btian 5 hours ago

$0.40

sidgtm 5 hours ago

In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot

aschobel 2 hours ago

Yah, I am pleasantly surprised at Grok Bot. Hopefully this improves CUA which has been a touch lacking w/ Grok 4.6. Grok 4.6 works but is slow compared to stuff like Astra Light.

guywithahat 5 hours ago

I've had really good experiences with Grok 4.6 and grok build. I've been playing around with tscircuit and it can write code with an understanding of spacial reasoning, while also importing cad components from different file formats into tsx, I've been having claude come in and try to error check it and so far claude hasn't found anything to improve in my three projects.

I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.

AM1010101 5 hours ago

Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?

ssutch3 5 hours ago

It did not. xhigh is new to grok.

forgot-my-pw 5 hours ago

Not sure on the API side, in Cursor you can always use 4.6 at xhigh.

ssutch3 4 hours ago

everfrustrated 3 hours ago

I think 4.6 got an xhigh after launch. The benchmarks seem to all have been against 4.6 high.

oh_no 3 hours ago

the AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.

sourcecodeplz 3 hours ago

looks like token efficient/verbosity took a big hit.

Output tokens from Intelligence Index:

- grok 4.6 (xhigh): 97M (for 44 score)

- grok 4.7 (xhigh): 240M (for 46 score)

oh_no 3 hours ago

which is crazy because this was grok's competitive advantage, worse than OpenAI models but better than everything else, now it's less efficient than Opus or Fable 5.1

jascha_eng 3 hours ago

32 on the omniscience index. Not terrible but far from Astra and fable: https://artificialanalysis.ai/evaluations/omniscience

andsoitis 5 hours ago

Congratulations to the team!

Invictus0 2 hours ago

SpaceX AI releasing "Grok" has to be some of the worst branding I've seen in my lifetime

DonHopkins 2 hours ago

xAI should offer persecuted White South Africans deep reverse-discrimination-victim discounts on Grok tokens.

MuffinFlavored 5 hours ago

If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low".

Is there a metric for like... time taken when comparing these two? I see score and cost.

If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?

Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".

inshard 3 hours ago

Any real world experience with Grok Ultra $300 monthly subscription vs Claude Code Max in terms of overall built work mileage, or general token limits?

mpalczewski an hour ago

Yeah I switched and the 300 plan is basically introductory 100/ month and I never hit the limit. While constantly hammering on it

gaigalas 3 hours ago

Pacing the frontier, with an aggressive release cadence. Gotta love the US tech industry.

brcmthrowaway 3 hours ago

Dumb question. Are these products really winner-take-all? Why is there such a furious rate of development?

dgellow 2 hours ago

It’s not at all winner takes all, it’s a race to the bottom. Models are becoming a commodity

hdhdjdif 3 hours ago

because boomers will give you free money + tip

musk can fund the space stuff with this

Saline9515 5 hours ago

I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

xmorse 5 hours ago

OMP is a joke. don't use that garbage

Saline9515 5 hours ago

Can you explain your opinion? I'm curious but such vague comments won't convince me.

samtheprogram 4 hours ago

xmorse 4 hours ago

polytely 4 hours ago

what do you use and why do you prefer it over omp

raincole 3 hours ago

unrvl22 4 hours ago

you are a joke if you think omp is a joke.

kristofferR 6 hours ago

What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?

Jcampuzano2 6 hours ago

https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

user43928 5 hours ago

> with a proposed shutoff date of November 12, 2026

That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

Jcampuzano2 5 hours ago

oh_no 3 hours ago

kristofferR 5 hours ago

That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

andsoitis 5 hours ago

mh- 2 hours ago

scottyah 5 hours ago

Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

kristofferR 5 hours ago

Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.

Iolaum 6 hours ago

I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.

Jcampuzano2 6 hours ago

https://openai.com/index/our-decision-on-cursor-following-it...

Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

babelfish 6 hours ago

ryeguy 5 hours ago

bluecalm 2 hours ago

Elon posted on X that Grok 4.7 is behind Claude and OpenAI for agentic coding:

https://x.com/elonmusk/status/2102082011233931762?s=20

so it's likely about usage in Cursor specifically.

DavCreator 2 hours ago

babelfish 6 hours ago

this is exactly it.

simianwords 6 hours ago

I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

Capricorn2481 6 hours ago

> The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

Razengan 5 hours ago

I want my sexbot to have a personality

nython 5 hours ago

artemonster 6 hours ago

I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea

xutopia 5 hours ago

Ask it to be critical of the birthday photos and see where that gets you.

artemonster 5 hours ago

raincole 3 hours ago

> The personality is bland

Sounds like a plus. Guess I will give Grok another try...

slowin 6 hours ago

This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.

ethagnawl 4 hours ago

> it doesn’t work nearly as hard

Until you ask it to start generating horrific imagery and then it's best in class.

Shiggy_ 2 hours ago

I would consider this a positive. I'm not interested in a company that wants to prevent you from using a tool that you're paying for.

The value of the internet is that people can share whatever they want, and use software how they want. This will mean that some people will abuse that. This is the tradeoff of a free society.

Tsarp 5 hours ago

Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model

forgot-my-pw 5 hours ago

It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...

myko an hour ago

After mechahitler and understanding how it got to that point I am embarrassed for anyone who pays attention to or uses Grok. Do better.

thih9 4 hours ago

I refuse to use Grok. Mostly because of the usual reasons - somehow this high profile AI model seems more disgusting than others and it is in a way impressive.

But also Xai doesn’t seem to care about user experience and long term support.

eknkc 4 hours ago

I am subscribed to ChatGPT, Claude, Kimi and GLM coding plans. 200$ one on GPT and the 20$ ish ones on all others. Recently added Grok and it has somehow bacome my second most used model.

For daily one off questions I prefer it because it is fast enough and I like the way it responds. I also use it for basic research like “find me a battery drill for this and that”.

Kimi and GLM feel extremely coding oriented. I use them for code reviews basically. I hate the way Anthropic models talk. GPT takes too much time and effort for that kind of stuff for some reason.

Grok happened to be a nice middle ground.

brandonagr2 4 hours ago

You should try it, it is less sycophantic than other models and is faster and better at most reasoning levels, don't confuse the twitter bots and services also named Grok with the frontier model itself

venzaspa 31 minutes ago

Perhaps he doesn't want to use it because it's owned by human being who many people view as vile.

swozey 4 hours ago

I can't take anyone seriously who uses grok seriously. I like to look at the cybertruck owners forum every so often because it's just... hilarious. And the amount of superfluous grok use over there is just insane. Half the posts I click in there will have a bunch of people dumping entire grok takes "why do people hate cybertruck owners?" "Because they're jealous and poor," sort of stuff that they just LOVE to post.

As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.

ElectronCharge 3 hours ago

Possibly interestingly, I can't take you seriously for having such a superficial approach.

You probably shouldn't cut off your nose to spite your face.

mempko 3 hours ago

enraged_camel 4 hours ago

This thing is DOA. They compared 4.7 xhigh to 4.6 high to make it look like it improved. The reality is pretty bad: https://x.com/chetaslua/status/2102087511367618942

outside1234 2 hours ago

Who uses this trash?

spiderice 2 hours ago

117 million people per month. Anything else I can google for you?

dom96 5 hours ago

It’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1].

1 - https://bench.killswitch-lang.org

sejje 5 hours ago

You'll have to include it in the future, or your benchmark won't be relevant.

For now, I doubt anyone would notice your protest if you didn't announce it.

peder 4 hours ago

I think you're seeing a big shift around it.... since it's been markedly cheaper and also still easily available from OpenCode, it's getting large enterprise traction.

mempko 3 hours ago

Yes, and that's a bad thing.

mempko 3 hours ago

Not sure why you are being downvoted. Until Musk owns up to his Nazi salute, I won't have anything to do with Grok, no matter how good or cheap it is. And yes, we need to keep talking about this because it's absurd.

BoumTAC 4 hours ago

Vals AI just affirm that Grok 4.7 is worse than Grok 4.6 (It ranks #24 on the Vals Index at 54.2%, down 5.0 points from Grok 4.6 (#14, 59.2%))

https://x.com/ValsAI/status/2102086608476590432

nostrebored 4 hours ago

Is this an ad for Vals AI? Looking at their website, the rankings don't mesh with my observed utility for almost any model outside of fable and astra being good-ish.

BoumTAC 4 hours ago

Absolutely not. I know Elon retweet them a lot when Grok is good. This is how I discover the company.

I like to follow them and look for benchmark for each LLM release.

nostrebored 2 hours ago

finnjohnsen2 2 hours ago

Is Grok relevant? Who uses it?

Maybe I'm in some kind of bouble but I have never met or talked to anyone who has used Grok.

Brendinooo 38 minutes ago

I bought a month of use for $100. Early impressions of Grok 4.6 was that it's good at talking about ideas (got me unstuck from a piece of writing that I was working on; that bought it a TON of goodwill) and just okay as a coding tool compared to my more extensive use of Fable and Opus. Not as smart as Fable but cheaper; not as capable as Opus 5 but way less annoying along the way. And it generates images, which Anthropic doesn't do.

Not sure if I'll hold the subscription but I could see myself working with it more.

cvwright 13 minutes ago

I tried it last month after hearing here that it didn’t speak in Claudisms.

As a chatbot it’s totally fine, virtually indistinguishable from Gemini or ChatGPT or Claude.

For coding it’s… okay. I tried 4.6 and it feels similar to Opus from 12 months ago, or maybe Sonnet from 9 months ago. YMMV.

chronogram an hour ago

I use it in the car to talk to because it's built-in, mostly for navigation via voice. It seems to be an old or quantised model, because it seems a few years old, and it has strict rate limits unlike Gemini, where it really just shuts off if you talk to it a few times in one day, but it's still fun to show passengers who are still new to LLMs.

dofm an hour ago

The only people I know of in the UK who will tell you that they use it are performatively alt-right or right wing. The kind of people who say "nanny-state" or "wokerati". GB News viewers. People who have an opinion on Meghan Markle that they think other people need to hear.

It's just an observation but so far a pretty solid correlation. Musk has so severely poisoned the well in terms of his UK reputation that the only people who are open about using Grok are... well, wankers is as good a word as any.

FWIW among the AI-using people, it mostly goes Claude Code, then Codex, then whatever runs on their Mac. The only Cursor user I knew has jumped ship to OpenCode.

moomoo11 40 minutes ago

losers

zug_zug 5 hours ago

Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no.

I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?

sejje 5 hours ago

If you haven't used it, how do you know if it's winning?

I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.

I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.

swalsh 4 hours ago

After the cursor aquisition it's become a quite capable coding model. If you take cost into account, it's close to the top. OpenAI is maybe still #1, but I'd put Grok at #2 (again, including cost as a factor).

grim_io 5 hours ago

It's probably the most aligned (to a single person) model out there!

puszczyk 5 hours ago

For me it works well for agentic coding tasks and terminal/unix/bash (in cursor and grok build); it's also token efficient and cheaper than gpt 5.6. It's def not as good as Fable for me (I haven't used Astra much, can't comment). So it's not the cheapest, not the most capable, but it has a good mix of it for my backend, go, infra work.

The voice is the same AI slop as the others imho.

(This is about Grok 4.6, I didn't test 4.7 yet).

edit: clarified I mean agentic coding tasks

svachalek 4 hours ago

The voice is the weird part. The early Grok 4 models had a very distinct presentation unlike anything else out there. Then suddenly it made a big jump in coding ability and started sounding just like every other model.