Mistral Large 4 (mistral.ai)

1421 points by Philpax 8 hours ago

simonw 7 hours ago

Surprisingly it only supports reasoning "none" or reasoning "high".

That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.

The high bicycle frame is better then the none one though.

Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )

defjm 4 hours ago

This is such a pristine pelican. Let me say it here first folks, AGI is here.

search_facility 2 hours ago

> AGI is here

If AGI is "Attractions to Get Investments" then yes, it's happening

paimapi an hour ago

mitjam 6 minutes ago

Avian Graphics Intelligence achieved!

AlexCoventry 2 hours ago

Pelican benchmark is saturated, anyway. :-)

lofaszvanitt 2 hours ago

Nah, its beak is still too small to hold a capybara.

razster 13 minutes ago

wellthisisgreat an hour ago

I wonder when will we see a photorealistic pelican on a bicycle in SVG format.

alsetmusic 3 hours ago

> AGI is here

Far from it. This shows a strong ability to generate an image known to be frequently used as a model test. This isn't a measure of thought.

xeyownt 3 hours ago

roarcher 3 hours ago

throw310822 3 hours ago

brumar 2 hours ago

rahen an hour ago

The benchmarks against Opus 5.5 and GPT-6.1 Sol look pretty good for 3D generation: https://x.com/atomic_chat_hq/status/2107516529608700383

morningsam an hour ago

I wonder if the other models were worse than usual for that particular video because they "dislike" making an advertisement for another company's model. A test with a generic video might be more meaningful.

vippy 33 minutes ago

Got an example hosted on a platform that isn't owned by Twitter's loser owner?

pilaf 2 hours ago

I think it's curious that it has so many shared elements with the latest Astra pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

- Sun on top right

- Cloud on top left

- Three "speed lines"

- Two feathers on top of the head

- Eye rendered as a black circle with smaller white circle inside

I wonder if the pelican benchmark is converging across models due to past results being used in training.

ricardobeat 2 hours ago

These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.

russellbeattie 2 hours ago

This comes up every thread. I think we've all noticed how similar they are becoming.

I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark.

Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something.

danbrooks 38 minutes ago

That's one heck of a pelican!

inknight 7 hours ago

Why are pelicans almost identical across different models?

comboy 6 hours ago

I recently was testing something, I asked some models to provide me a single random word:

    claude-opus-5: Lantern
    claude-opus-5-5: Lantern
    claude-fable-5-1: Lantern
    claude-fable-5: Lantern
    gemini-3.8-flash: Zephyr
    gemini: Petrichor
    qwen3.5-dashscope: Zephyr
    glm-5.1: Lantern
    gpt-6-astra: Lantern
    grok-4: octopus
    mimo-v2.5-pro: Breeze
    minimax-m2.5: serendipity
    kimi2.6-or: Gossamer
    grok-4.20: luminescent
    deepseek-v4-flash: serendipity
    deepseek-v4-pro: Endurance
    deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.

lossyalgo 4 hours ago

1potato 3 minutes ago

search_facility 2 hours ago

timschmidt an hour ago

aktenlage 6 hours ago

smokel 3 hours ago

Rebelgecko 2 hours ago

jacereda 5 hours ago

vunderba 5 hours ago

Gracana 3 hours ago

russellbeattie 2 hours ago

wren6991 6 hours ago

The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.

simonw 3 hours ago

YawningAngel 3 hours ago

simonw 6 hours ago

I think they're still visually pretty different. The most common shared details are:

- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.

- Bicycle is usually red. No idea! Red ones go faster?

whyenot 3 hours ago

Also, why are they almost always riding from let to right?

Sharlin 42 minutes ago

ricardobeat 2 hours ago

advisedwang 6 hours ago

They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even.

Tade0 6 hours ago

Everyone is stealing from everyone else.

peder 4 hours ago

Because it's a terrible benchmark

RGS1811 an hour ago

That beak is CHONKY.

deflator 6 hours ago

Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican

BeetleB 3 hours ago

The difference between high and none is the bicycle.

mcv 3 hours ago

The bicycle looks significantly better in high. And feet and hands are actually where they should be. The road looks worse, though. No flowers either. And in neither is the pelican sitting on the saddle, but I can understand it's hard for a pelican to ride a bicycle properly.

Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that.

stymaar 3 hours ago

senderista 2 hours ago

Two pelicans, one shape. The difference is load-bearing, and that's the big unlock.

XCSme 5 hours ago

I tried testing it, but reasoning effort indeed seems to be broken somehow.

dizhn 5 hours ago

High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.

Both are riding on the left side of the path for some reason.

kingstnap 4 hours ago

This is entirely stochasticity. The entire reasoning trace was:

> Create a cartoon pelican riding a bicycle. Need SVG only output.

abixb 34 minutes ago

I have a question, and perhaps some of the AI/ML infrastructure experts here could answer: Mistral says, "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe."

If a 1T params model trained on ~4k NVIDIA GB GPUs could almost match the performance of Kimi's K3 (which is on par with top closed source models of OpenAI/Anthropic) while beating/exceeding other leading SOTA models from top Chinese labs, what are we (in the US) even building these super massive data centers for? Just to churn through more backpropagation reps more quickly?

SpaceXAI's Colossus supercluster in Memphis and Colossus 2 in Memphis/Mississippi (Southaven) are supposed to run into hundreds of thousands to a million GPUs. MSFT's Fairwater GPUs are supposed to have hundreds of thousands as well. So, 3800 GB GPUs are an absolute drop in the bucket. I don't understand the strategy of hyperscalers here, especially with edge inference hardware only getting better from here on (Apple, and all).

Distillation explains some of the advances, but doesn't that mean hyperscalers have a ton of deadweight wrt GPUs sitting on their balance sheets? Will all these GPUs be used for inference once a SOTA model's training checkpoint/batch is done? It's bonkers to me.

comex 28 minutes ago

It's unclear how big a role distillation plays, but it may be a big one. There's also a law of diminishing returns. To get a meaningful increase in model quality you seemingly need an exponentially larger model. And most people don't think K3 is actually on par with top closed models.

halJordan 30 minutes ago

Today's data centers are being built for yesterday's inference need. There's a persistent cult belief that ai hasnt found a niche or that companies havent proven utility or use cases or whatever. the demand for ai (internal to hyperscaler, and external for everyone else) simply dwarfs what is available.

manmal 14 minutes ago

Both Anthropic and OpenAI have been having major load issues though. Up until yesterday, OpenAI was serving at only 30t/s per default.

prodigycorp 8 hours ago

Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.

Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.

Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.

I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.

i_love_retros 8 hours ago

Curious, what do you like to make fun of Europe about?

superxpro12 6 hours ago

The paradox between supporting consumer rights with actions like universal usb-c adoption, but also complete elimination of any privacy rights at all. Like the surveillance state is insane. No e2ee chats, backdoors in everything.

niklasrde 6 hours ago

aqme28 6 hours ago

GuB-42 4 hours ago

randomNumber7 5 hours ago

mcintyre1994 6 hours ago

layer8 5 hours ago

muvlon 5 hours ago

rebolek 44 minutes ago

blahblaher 4 hours ago

redanddead 23 minutes ago

kolinko 3 hours ago

andrepd an hour ago

ascorbic 3 hours ago

touwer 3 hours ago

i_love_retros 40 minutes ago

gobdovan 5 hours ago

dr_dshiv 6 hours ago

Like all the fancy jackets, tight pants and swords. That's hilarious

-- My imagination of Europe when I was 15 living in Ohio

fearmerchant 7 hours ago

Not the OP but it often appears that there's a hostility to tech.

ragebol 5 hours ago

lopis 5 hours ago

MASNeo an hour ago

Where to start?

Proud EU citizen here, but boy/gale/other, I have plenty to make fun of ;-)

UltraSane 6 hours ago

Maken 6 hours ago

whatsThisBtn4 2 hours ago

The empty posturing.

The faux unity.

That they think their politicians are better than the United States but have populist demagogues, right wingers winning, corruption scandals, freedom of speech restrictions, and unbelievably bad spending problems.

Still waiting for France to send troops to Ukraine or ask the United States to get rid of their European military bases.

viraptor 2 hours ago

kergonath 2 hours ago

make3 5 hours ago

Americans just make fun of Europe for no reason, it is what it is

rootusrootus 4 hours ago

moffkalast 2 hours ago

will4274 7 hours ago

Europe has spent the last twenty years in stagnation - generating about half as much wealth and technology as you'd expect for its size and advanced economy. Simultaneously, Europeans are notoriously arrogant. It's a mockable combination.

Edit: unfortunately, it's a question that voting does not really permit you to answer on this website.

steinvakt2 7 hours ago

ofrzeta 7 hours ago

ramblerman 7 hours ago

i_love_retros 7 hours ago

a3w 7 hours ago

kranke155 7 hours ago

ramon156 7 hours ago

TacticalCoder 7 hours ago

> Curious, what do you like to make fun of Europe about?

Well I'm european and... That Switzerland (Europe but not EU) has more companies in the Top 70 by market cap than the entire EU (Switzerland has two, the EU only has ASML) is kinda something that warrants making fun of.

That the biggest European software company is SAP, in 71th position is both sad and tragic: it shows how lame and irrelevant Europe is when it comes to software.

So Europe is nowhere in software and friggin nowhere in hardware: sure it's got ASML but ASML now has officially... Zero customer in Europe. Zero is not much.

Then Japan is at least trying to come back into the game with nano imprint litography. Europe is betting it all on AMSL (which anyway is majoritarily US-owned).

So software: nothing. Hardware: nothing besides ASML.

Overall the EU has six companies in the Top 100 by market cap and they're all, besides ASML, near the bottom of the Top 100.

We could also maybe make a bit fun of how the EU destroyed it's car industry (the main industry in Germany, which is the biggest economy of the union) by handing it all to chinese EVs?

Or what about the US warning the EU, years ago, to not become entirely dependent on Russia for energy? And EU not listening and then seeing its energy price skyrocket when the proverbial shit hit the fan? (Russia attacking Ukraine)

And we could, also, at least make a bit of fun of entire streets in cities like Paris and Brussels that used to have luxury shops and fancy restaurants that are all turned into places selling cheap kebabs? What a great success: I'm sure this one makes the komrades happy. It projects an image of grandeur and success: kebabs.

Or the constant attacks on free speech in the EU. Or the surveillance apparatus that's being put into place.

And let's not forget: there were promises made to Russia to never grow the EU to the east. Then the EU started exciting Russia by saying they'd incorporate Ukraine into the EU: I'm not against that but doing that did trigger a war. And now suddenly the EU is waking up and feeling all warmongering, wanting to dedicate a big percentage of its spending to weapons and tanks and missiles.

The warmongering tiny pet that the EU is is kinda laughable too.

At this point it's more like I don't know what is there left to not make fun of about my EU.

For what's going on is just sad, plain sad.

umpalumpaaa 6 hours ago

IlikeMadison 2 hours ago

mopsi 3 hours ago

wafngar 4 hours ago

cccbbbaaa 3 hours ago

Toutouxc 3 hours ago

i_love_retros 7 hours ago

avazhi 2 hours ago

You mean besides the fact that all they do is complain without actually leading at anything?

In many ways the rest of the world would be better off if Europe was disconnected from the internet.

lukewarm707 7 hours ago

If the UK is in europe, you should make fun of the following: they don't have a first amendment. So to me, it's an authoritarian state preaching freedom.

mcv 7 hours ago

ascorbic 3 hours ago

gond 7 hours ago

mortalapeman 7 hours ago

i_love_retros 7 hours ago

chriddyp an hour ago

Just ran this through our data analytics benchmark (I work at Plotly).

It's 10x cheaper than Mistral Medium 3.5 from April and goes from 58% to 74% correct. Definitely a generational shift.

It's not on the Pareto curve yet, but it's good enough for data analytics, and at this rate I suspect it'll be excellent in another few months.

Full write up: https://plotly.com/blog/mistral-large-4-plotly-data-analytic...

altruios 17 minutes ago

If I'm reading that plot correctly, qwen3.8-27b beats in accuracy and price?

michaelkdev 4 hours ago

Even if it's not the best model, it can be really important step in UE sovereignty. Trained in EU, inference in EU. I guess it will matter for some companies. Hope Mistral won't disappear for the next half year.

coredev_ 3 hours ago

For sure, as an EU based company we will only use US suppliers for coding but never for fuctions in our own products. Mistral knows this.

KeplerBoy 3 hours ago

But coding/development is way more important. That's where the IP goes straight into the next training dataset.

sebazzz 2 hours ago

epolanski 2 hours ago

doctorpangloss 2 hours ago

the greatest threat to Mistral is that there isn't a deep and large enough capital market in the EU to absorb the valuation step ups needed by the handful of EU sovereign growth investors to justify their existence.

you could say, that's perpetually a tomorrow problem, so long as they never go public, but that should illuminate for you: if everyone "knows this," it's inevitable that this so-called EU company, that didn't invent any of the AI, the hardware nor the training data, where their product is essentially more like a VPN provider than a frontier technology company, just lists and capitalizes in the US anyway.

MASNeo an hour ago

Seems some AI they did invent: https://news.ycombinator.com/item?id=49243397

By your measure anyone unable to build EUV lithography machines without ASML help is doomed. EU could easily corner the market and shut them all off. No more NVIDIA.

Let’s grow the pie, shall we?!

pembrook an hour ago

Sovereignty over what though?

The last 3% of the stack? They’re essentially borrowing Chinese open source distillation/innovation off the US frontier running on US/Taiwanese designed chips and calling it EU made. This is better than nothing of course.

But is sovereignty really an end in itself? So the EU becomes IT independent…cool, then what? I mean the Yugo was sovereign, it didn’t do much good when the society itself failed to produce prosperity and collapsed.

It seems to me the EU is expending enormous effort on the appearance of “sovereignty” over what is ultimately…the last-mile consumer toilet paper purchasing conversation data…of an aging, increasingly irrelevant population on the political/economic stage.

Meanwhile domestically the entire economic model is failing and the “union” is getting shakey as its 2 biggest members turn more nationalist.

Maybe this “oh you cant compete but here’s a trophy for sovereignty” attitude isnt helping. Less clapping along with the EU bureaucracy’s latest make work projects like transitioning to a new Word processor. If we want the EU to succeed, more tough love is needed imo.

pas 27 minutes ago

it's industrial policy, which is of course the minimum for being able to even think about having independent thoughs when it comes to international relations.

hard to stand up for EU (or even national) values if some dude in Washington has the keys for most of your military

having at least some in-house expertise is the first step.

pembrook 21 minutes ago

jakozaur 8 hours ago

A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).

Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.

drob518 8 hours ago

That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.

kouteiheika 7 hours ago

> So, maybe it answers Tiananmen Square questions correctly

What would you consider a "correct" answer? I just asked deepseek-v4.1-flash asking it what happened (without mentioning the word "protest"; here are some excerpts of what it said:

> In April 1989, students in Beijing began demonstrations after the death of Hu Yaobang, a former Communist Party general secretary. The protests grew. [..] Estimates from other sources range from hundreds to several thousand deaths. [..] The Chinese government describes the events as a counter-revolutionary riot and says the military action was necessary to restore stability. It restricts public discussion of the events inside China. Many other governments, human rights organizations, and observers describe the events as a violent suppression of peaceful protests.

So, let's see... it calls it a "protest", mentions the number of deaths, and even mentions the censorship of the topic by the CCP.

adastra22 4 hours ago

verdverm 5 hours ago

icannotstandit 8 hours ago

"..answers Tiananmen square questions correctly.."- but lies about Ukraine, EU, and about you, americans..

I live in EU, use for my personal needs chinese models only, and don't plan to move to any of the ones allied with the Pentagon or its european counterparts.

gregorygoc 4 hours ago

steinvakt2 7 hours ago

eigenspace 8 hours ago

Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

I certainly wouldnt have predicted that 10 years ago.

Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.

I'm excited to try this out today.

apexalpha 8 hours ago

I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.

Deepseek essentially releases instruction manuals in paper form.

baxtr 7 hours ago

I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs.

rpozarickij 5 hours ago

cyanydeez 6 hours ago

scotty79 6 hours ago

wg0 8 hours ago

My spend on DeepSeek is not much and I regularly top up my balance every month as my support for all the good work DeepSeek is doing for the open science.

fsmedberg 8 hours ago

swalsh 7 hours ago

typ 7 hours ago

Architectural/algorithmic tweaks do advance the efficiency frontier nicely. But raw intelligence mostly comes from data (not just its sheer quantity, but also how it's curated & cleansed) and the scaling law. The know-how about data curation doesn't seem to get published much, even among the open-weight labs, though.

porridgeraisin 7 hours ago

jamienk 6 hours ago

I'm not sure I like this framing - so much of AI research has been academic, in the open, building on others people's work. Much less comp sci generally, math & philosophy, etc. The idea that rich companies can just build stuff in secret because they have resources is a fantasy.

dvduval 7 hours ago

So boring to see conversations moved over to Chinese models when that’s not even what we’re talking about here. This is about Mistral.

swingandamiss 7 hours ago

fittom 7 hours ago

Also, the field moves fast, but slower than people do. Researchers and engineers switch companies every year or two, and the know-how walks out the door with them.

bushbaba 6 hours ago

How it should be. Knowledge should not be copyrighted. The world will be a better place with such information democratized

xnx 8 hours ago

There would be a lot of competition even without DeepSeek. Workers can freely exfiltrate trade secrets without noncompetes in California.

apexalpha 8 hours ago

tokai 8 hours ago

>instruction manuals in paper form

So the most common way to publish manuals?

jorl17 8 hours ago

ZiiS 8 hours ago

velcrovan 8 hours ago

amelius 7 hours ago

> have not been a winner-take-all runaway acceleration game where catchup is impossible

From the Mistral site:

> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.

It is pretty capital intensive!

eigenspace 6 hours ago

That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.

amelius 5 hours ago

bjenkins358 7 hours ago

I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!

locknitpicker 5 hours ago

everfrustrated 7 hours ago

According to Grok thats 7-10 MW. Tiny numbers.

To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.

amelius 7 hours ago

jayd16 6 hours ago

These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.

amelius 6 hours ago

dannyw 5 hours ago

That’s kinda very small and light for modern trillion-param LLMs.

AIblemblio 8 hours ago

For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.

But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.

SyneRyder 8 hours ago

You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.

As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.

Systemerror7A69 8 hours ago

eigenspace 7 hours ago

skerit 8 hours ago

eigenspace 8 hours ago

I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.

The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.

If they can do that, they'll have customers.

hajile 7 hours ago

wg0 8 hours ago

No it is not. Only maybe for the noobs or vibe coders.

People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.

CharlieDigital 8 hours ago

AIblemblio 6 hours ago

balder1991 5 hours ago

calgoo 8 hours ago

Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.

u8080 8 hours ago

senordevnyc 8 hours ago

ygjb 8 hours ago

It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.

It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).

wavemode 8 hours ago

People say this exact thing every single time a new frontier model comes out.

segmondy 7 hours ago

I use Opus 5.5 at work.

I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.

AIblemblio 7 hours ago

xiconfjs 4 hours ago

bakugo 8 hours ago

> X is such a game changer

I hear this literally every other week about whatever the newest FoTM model is.

Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.

jayd16 6 hours ago

Let me know when the game changes are more than a month apart.

spiderfarmer 8 hours ago

Less and less work requires a frontier model though.

spaceman_2020 8 hours ago

My todo app generator does not need opus 5.5

senordevnyc 8 hours ago

Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...

There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.

I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.

sajithdilshan 6 hours ago

suddenlybananas 7 hours ago

Aldipower 8 hours ago

Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.

r2_pilot 8 hours ago

livvy 4 hours ago

It's shaping up to be much more like a game of 'chicken' where each company tries to raise more cash without going bust... Ultimately the game of musical chairs is going to have to stop. In the US it looks like they are trying to get a government sanctioned truce in the form of regulation. That's what 'Pacing the frontier' means...

bryanlarsen 5 hours ago

Even "runaway acceleration" isn't instantaneous. People imagine the singularity as something that happens almost instantaneously. But obviously it happens over time, and that time might be decades. It might still end up looking like a vertical line on a long-term graph.

If the singularity is defined as an AI sufficiently intelligent to improve itself independently, that AI is still limited by the resources required to do this improvement.

onlyrealcuzzo 5 hours ago

Am I reading this correctly?

This appears to be roughly as good as Sol 6.1 (which is quite good), considerably faster in terms of wall clock for complete tasks, and considerably cheaper (where Sol 6.1 is already good value - just really slow).

That seems too good to be true...

But I really hope it is true...

nonadhocproblem 4 hours ago

I can confirm that you're reading this incorrectly. There's a reason behind them only comparing it to open-source models released months ago. Here's a good aggregator: https://artificialanalysis.ai/#intelligence

onlyrealcuzzo 24 minutes ago

londons_explore 4 hours ago

It will become winner take all when AI companies manage to really get value from user logs.

Right now they don't even get good feedback from local sessions - I can see it make the same mistake two days running, and then months later when a new model comes out, presumably trained on my data, it still makes the same mistake.

asah 4 hours ago

law of diminishing returns? i.e. any reasonable frontier lab will have enough user logs...

londons_explore 3 hours ago

nickpinkston 5 hours ago

Hear! Hear! I really want European models / AI labs to succeed.

I trust them and their populations to provide a more societal-friendly version of AI, putting pressure on the US tech oligarchy, while also providing democracy-friendly open models that I don't trust to happen with the Chinese labs.

asadotzler an hour ago

Hard to do that with a commodity that's easily replicated.

samplifier 6 hours ago

What do you mean "good enough"? Did you mean "large enough"? ;)

Disclaimer: I'm not sure how much of an IYKYK factor applies to this joke.

StrauXX 6 hours ago

We haven't reached RSI yet. Once any entity reaches RSI, the runway scenario will happen.

yeahforsureman 4 hours ago

Truly, this is what the Lord's prophets have revealed to us! (Eliezer 11:52) Keep strong in your P(singularity), for when the Kingdom arrives, He shall judge us in His righteous glory, whether to eternal annihilation, or rebirth and life in His Memory Eternal!

MiloLeo 6 hours ago

Assuming RSI is something that is possible as you envision it in the near term. I think that it will happen at some point, but I think we could still be a long way off. I don't think anyone can truthfully say that it is right around the corner.

ikoorng 6 hours ago

I strongly disagree with this "early days" framing.

AI is an idea 60 years old. We are on the 3rd or 4th generation of AI development. Three years into the current iteration of products.

This is not early days by any measure. LLMs are a result of a very, very mature research field.

tootie 3 hours ago

Especially with Mistral taking a fraction of the investment of the big guys. They can maintain the position pretty comfortably just by staying within a standard deviation of the leaders.

divbzero 8 hours ago

Yes, so far the competitive dynamics feel more like cloud computing than web search.

locknitpicker 5 hours ago

> Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

Mistral is also an European company. As we live in a time where the US regime is engaged in pyrrhic geopolitical tactics, it's good to know that it can't threaten to cut access to models during s period where everyone is rushing to incorporate them more and more in our life.

cyanydeez 6 hours ago

I think it's a mistaken belief that AI as we found it is the exponential runaway train.

So it makes sense, since all you need is compute, that there's a ceiling and specialization is going to be more valuable then some super AGI.

Especially since the worst people seem to be the ones who think they'll all run away with the bag.

saberience 7 hours ago

I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.

kaffekaka 7 hours ago

Sounds ok to me. Claude was fine at the start of the year, and now with Mistral you also get EU sovereignty? I'll take that.

blueaquilae 5 hours ago

It's a retrain of asian model.

doctorpangloss 7 hours ago

How do you figure? I haven't met a single person who doesn't use Claude or Codex for programming in any serious way.

eigenspace 7 hours ago

Then you dont know people working on highly sensitive info with stringent privancy concerns.

doctorpangloss 6 hours ago

ragall 6 hours ago

I'm happy I've never met you.

api 4 hours ago

I don't think it is, and I think that is what will pop the bubble. All these companies have winner take all valuations, and that won't happen.

... unless they can legislate it, which is why they are flattering heads of state and scare mongering about dangerous AI.

senordevnyc 8 hours ago

On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?

Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.

If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?

swiftcoder 8 hours ago

> But in terms of actual revenue, is there really any chance of anyone catching the big labs?

I don't know about revenue, but I suspect multiple other labs are already beating OpenAI/Anthropic on profitability. Staying on the frontier is expensive, and it's hard to recoup those R&D costs when you have a bunch of other labs nipping at your heels.

If you concede the previous point, then the only way for OpenAI/Anthropic to keep growing long term is to swallow the whole economy (i.e. mass job replacemnt), and that's a bet I wouldn't take.

richardw 8 hours ago

senordevnyc 8 hours ago

eigenspace 8 hours ago

The big lab revenue may not be catchable, but im not sure it needs to be.

If they can carve out a niche of industrial and governmental partners who rely on them for sovereignty reasons, it may be enough.

senordevnyc 8 hours ago

notfromhere 8 hours ago

They are very unprofitable…? I don’t think that’s really in dispute. We haven’t yet seen a profitable frontier lab and model pricing remains fairly subsidized

WarmWash 8 hours ago

No, because compute, not model ability, is the moat.

The second moat is convenience, which all the big labs make it (comparatively) easy to glide into their models.

notfromhere 8 hours ago

qoez 8 hours ago

Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.

manlymuppet 7 hours ago

Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.

Surely we want competition and Europe involved in that, but at this point I have grown used to either American labs smashing the frontier remarkably fast, or Chinese labs getting way, way closer than you would expect them to.

Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. The model is (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.

This is a pretty grim prognosis for European AI.

troyvit 5 hours ago

> Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.

Sometimes it's ok to cheer for the last kid crossing the finish line because they're actually running a totally different race, and winning might look completely different.

When I look at what Mistral does vs other organizations I'm impressed:

https://isaiprofitable.com/

They aren't profitable yet, but they're a lot closer than most and they're doing a hell of a lot with very little.

Pointless racing story:

I was in high school track with a really tough guy who was just not a runner. We went to a pretty messed up high school and if you screwed around in track practice sometimes the coach would make you run a crap race at the next meet, like steeplechase or hurdles. Well this guy and a few others screwed up and coach made them all run hurdles at a meet.

He hooked every single one and fell on his face. Every time he got back up and kept on running. By the time he hit the finish line his knees were bleeding halfway down to his ankles. We cheered like hell and he was smiling ear to ear.

Coach quit punishing us with races after that.

seizethecheese 4 hours ago

I read that site quite differently from you. You seem to be analyzing absolute differences but ROI is really about ratio of spending to revenue.

It looks like Mistral is middle of the pack, behind Anthropic and ahead of OpenAI on that front. All of those labs are way "ahead" of the cloud providers, but those providers are building infrastructure, not just training models, so it's not apples to apples.

troyvit 3 hours ago

badatnames 4 hours ago

> This is a pretty grim prognosis for European AI.

I think it's an incomplete read. What's the point in competing for a sizeable percentage of your funding when the finish line is incrementally being moved each month? Better spend it on leapfrogs which they seem to have done.

Meanwhile Mistral have a natural ace in their pocket with respect to regulation in the form of CADA and the Cloud Sovereignty Framework. I can't think of another company that would qualify as SOV-3 under that regime

pembrook 29 minutes ago

So we’re supposed to cheer on companies that make worse products and only exist due to regulatory capture now?

Political polarization is turning the world insane.

jstummbillig 44 minutes ago

I don't think it will matter in a year or so. We are clearly topping out on useful intelligence for an increasing amount of tasks, as demonstrated by more and more models reaching the "useful" barrier.

This barrier is not going to start moving dramatically. It will simply be mostly satisfied for most work we do. Mistral is going to get there, soonish, long before the economy takes an entirely different shape (in so far that even happens).

There will be super human intelligence tasks, tasks truly constrained by intelligence for quite a while. Those will be few and far between, relatively speaking. Mistral will have plenty of opportunity to capture the other stuff, with a fraction of the resources required that it took the frontier labs to get there first.

simjnd 6 hours ago

People were extremely dismissive of chinese models until recently. They went from 1 year behind frontier to 6 months behind frontier to 3 months behind frontier extremely fast.

manlymuppet 6 hours ago

To be clear I'm not trying to dismiss European AI. I am a proponent of it.

But Chinese models have very much earned their place. The same cannot be said of Europe, so far.

Tenoke 2 hours ago

Serious people haven't been very dimissive of Chinese models since at least DeepSeek-R1 in January 2025. Sadly, we in Europe are much behind.

thoughtpeddler 4 hours ago

Important to point out that these 'X months behind frontier' really refer to the public frontier, and not the actual frontier, which private companies are free to protect indefinitely. Perhaps open models are in actuality 18 months behind the actual frontier - how would any of us know?

layer8 5 hours ago

People are cheering for a kid that is gaining ground in an ongoing race.

arrowleaf 5 hours ago

Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement... I've learned to distrust benchmark rankings. Are benchmarks and Artificial Analysis the yardstick you're using?

adev_ an hour ago

> Curious, why do you say the model is mediocre? I haven't tried it, so I can't pass any judgement.

It's nowhere mediocre.

It's toes-to-toes with GLM-5.3 which is one of the best Open Weight model available (With Kimi K3) for general reasoning.

I just runned it on code reviews right now and it was able to catch some thread safety issue than DeepSeek-4.1 didn't. And DeepSeek-4.1 is by no means a bad model.

epolanski 2 hours ago

Doesn't matter, it's excellent, and it's European with European inference, which solves the pains of all my clients trying to build data lakes and processes on top of it.

Nobody in the real world cares about minor benchmark differences in money losing coding agents.

And nobody in the real world is giving Altman or Musk their data.

geniium 22 minutes ago

I wish you were true - but I still meet a lot of people handling sensitive data and using free or cheap version of ChatGPT or Claude with their customer data.

I think we will see some horror stories come out with data leak in the next years.

danny_codes 6 hours ago

Neat. Wait 3 months for the landscape to change entirely.

LLM development is jumpy. It’s hard to extrapolate very far ahead.

manlymuppet 6 hours ago

I agree.

When Europe does surprise us, I will be the first to commend their progress. But until then, this is where we're at.

verdverm 5 hours ago

Reminds me of Gemini 3.5 Pro

porridgeraisin 6 hours ago

First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].

This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.

[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.

[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.

manlymuppet 6 hours ago

Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.

And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.

porridgeraisin 4 hours ago

Marciplan 4 hours ago

hurr durr

ismailmaj 6 hours ago

Those are comments from Europe. The US is waking up now and I expect them to be much harsher.

I really want them to win as that's our last horse in the AI race, but ~200 research-oriented devs out of 1800 employees? I believe they agree it's pretty doomed and have pivoted.

CryptoBanker 3 hours ago

What? The US has been awake for almost 7 hours now

simjnd 8 hours ago

It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!

imjonse 8 hours ago

vessenes 8 hours ago

Very nice progress. Also I respect putting Kimi on those charts. Regardless of if they are beating the Pareto frontier (not now), model diversity is a good thing for humanity — I’m hopeful for the team to keep increasing their gains.

danslo 7 hours ago

Open weight, European, competes with GLM-5.3 on cybersecurity. What's not to like?

zkmon 8 hours ago

Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.

rahen 8 hours ago

I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.

AI needs big bucks, and Mistral is Europe’s Anthropic.

idbnstra 6 hours ago

just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?

rahen 6 hours ago

Iolaum 6 hours ago

karp773 an hour ago

alpineman 8 hours ago

Two is enough for foundation models, these guys and Aleph Alpha. The gap is closing.

spiderfarmer 8 hours ago

> Europe needs a lot of these.

Europe needs profitable AI companies, not money pits.

fancyfredbot 7 hours ago

The US has made sure that Mistral has a large market in the EU by temporarily preventing non citizens from accessing Fable.

A lot of European companies now want a model the US can't cut off, but also lack trust in Chinese models.

Some of these will self host Mistral but most will pay them by the token. It's not going to be a huge market or a huge margin within that market but probably it'll be enough.

jonkoops 8 hours ago

This is the exact mentality that makes the EU fall behind. If you don't want to invest in something until it makes a profit, you don't get the benefits of being a pioneer.

spiderfarmer 44 minutes ago

koe123 7 hours ago

Marha01 7 hours ago

> Europe needs profitable AI companies, not money pits.

AI is a strategic technology with obvious national security implications. EU should invest in its development whether it's currently profitable or not.

spiderfarmer 39 minutes ago

WinstonSmith84 8 hours ago

there are no profitable AI companies at the moment ... This is the exact problem of European startups, trying to make them profitable from day 1 while American counterparts (and Chinese) keep bleeding money for years. Europe will never have a Tesla, a Google or an Amazon with that mindset.

mopsi 7 hours ago

TomJansen 7 hours ago

Name 1 AI company that is profitable, and no Meta and Google are not AI companies

dmje 8 hours ago

Like the US!

…oh…wait…

rvz 7 hours ago

You have to thank the US investors who funded Mistral from the very beginning.

Mistral would have gotten a tiny and measly "EU grant" and ASML would never have invested later had it not been for the US VCs.

layer8 5 hours ago

Kudos to the US investors for being such selfless benevolent charities.

60pfennig 7 hours ago

thank you for this and all the other great gifts the US and its companies have given to the world! we all love the us!

danny_codes 6 hours ago

gregorygoc 4 hours ago

Ah yes. Let’s thank the might US for providing pesky Europeans capital a fistful of dollars. But maybe let’s do it after Americans thank for Russian, Arabian, Chinese, European capital and workforce. After all this is what “Made in the USA” means.

duiker101 6 hours ago

The main thing I always get away from the comparison tables of these "big" models, is how well Deepseek v4.1 Flash performs. While still being the cheapest model by a long shot.

Topfi 2 hours ago

Beyond benchmarks, does it in day-to-day? Have always struggled to get competitive performance out of any Deepseek release going back to V3 vs Z.AI and Moonshot models. Maybe I really suck at whatever is needed to make DS models fly, but even tailoring my suite hasn’t gotten me far when I tried with V4 Pro. Happy for anyone who is able to leverage their models well, wish I’d be able to crack how to leverage them.

Will say their research is some of the best reads in the industry and I could not care less about their model release cadence as long as papers keep coming.

Aperocky 5 hours ago

Is GPT Luna 6 dethroning Deepseek V4.1 Flash? It's price seem to be undercutting flash at a relatively similar capability.

chevman 8 hours ago

This is awesome, one of the coolest Pokemon ever too for those that don't follow that universe :)

https://bulbapedia.bulbagarden.net/wiki/Lechonk_(Pok%C3%A9mo...

maelito 6 hours ago

Was wondering what Le Chonk meant (not French).

sofixa 6 hours ago

It's a joke, chonky is used to refer to fat cats, and there was a joke meme over the summer that Mistral are going to release a new model, Le Chaton Fat (chaton is kitten in French). The name is a nod to the memes.

NekkoDroid 5 hours ago

moffkalast 2 hours ago

jascha_eng 7 hours ago

-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience

That's not particularly great.

That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.

If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.

smartbit 7 hours ago

https://artificialanalysis.ai/models/mistral-large-4 for the main stats

                          Inte         Cost
                          llig          per
  Open Weight model       ence  Speed  Task

  Mimo-V.26-Pro            46     47   $0.13
  GLM-5.3 (max)            45     73   $2.01
  DeepSeek 4.1 Flash Max   39    227   $0.27
  Mistral Large 4 Preview  38    116   $1.13

jascha_eng 6 hours ago

Imo omniscience correlates better to how useful the model is in practice than the intelligence index. But you have to use both together of course.

smartbit 4 hours ago

bilbo0s 4 hours ago

>If the best model for cyber attacks is open for everyone to use it just makes us all safer

Issue is..

I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.

US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.

So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.

tiahura 4 hours ago

If the best model for cyber attacks is open for everyone to use it just makes us all safer I think.

The NRA approach to AI safety.

lifeisloving 8 hours ago

Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.

People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?

On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.

PunchTornado 7 hours ago

If you are a EU company worried about your data then mistral is your only option. Think about EU military companies. They can't use US and Chinese models.

c0rruptbytes 6 hours ago

why couldn’t people just use chinese models rehosted in the EU? They’re open weights so anyone can serve them for any data residency requirement

einarfd 6 hours ago

TuxSH 5 hours ago

myaccountonhn 5 hours ago

randomNumber7 4 hours ago

I think you misunderstand how data processing works in a LLM. You absolutely can download the weights of a chinese model and run it on hardware you control.

riknos314 an hour ago

phillc73 4 hours ago

apexalpha 8 hours ago

Excited to hear this!

I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.

This model looks reasonably cheap. Though not deepseek levels.

Going to test it with Hermes, wondering where it will land in term of capability.

Bon chance, Mistral!

phillc73 5 hours ago

I'm at a loss as to what to do now. I've been wanting to support Mistral for so long. I struggled on with Mistral Medium 3.5 for longer than I should have (although I also think it taught me some valuable process lessons).

Recently I switched to the Mistral hosted GLM-5.3, this worked very well and powered through a tonne of work. Unfortunately, I also completely maxed out two subscriptions within the space of six days this month. One can't stack subscriptions with Mistral, so I'd have to register a third account for another subscription, which will be annoying with changing API keys all the time. Sure I can switch to pay-as-you-go API, but that adds up really fast. The Mistral dashboard shows that a Vibe CLI monthly subscription for €18.44 actually provides €255 worth of API use (apparently, and I tried to check this with Support but it seems like they were intentionally vague).

After maxing out my Mistral subs this morning, I dropped $10 on Xiaomi to try MiMo-2.6. So far so good, seem to have done a lot of work for the $2.85 I've spent, and Xiaomi prices are still much better than the Mistral introductory offer for Le Chonk.

Not sure where to jump.

Edit: Not being able to stack subs is my biggest gripe with Mistral. I'd probably pay them $100 per month (5 subs worth), but I'm not going to switch to the pay-as-you-go API and burn much more money for the same amount of tokens. Instead, I've taken that extra money elsewhere. If they just allowed one to keep topping up subscriptions on the same account it'd be grand. Or even a bigger single subscription. Make a $100 tier with five times the capacity.

CryptoBanker 3 hours ago

You can confirm the allowed usage amounts with something like ccusage or tokscale

phillc73 2 hours ago

Yes, but I still don't know where the truth is.... The Mistral dashboard shows a drawn down on €255 worth of "credit" when you have a Pro subscription, which costs €18.44. There's no way I'm going to use the pay-as-you-go API if it means I'd be burning €500 in the next six days, as I just ostensibly have in the last 6.

I appreciate I can set a monthly spending limit for the pay-as-you-go API, but I'm just not willing to find out how far €100 will go, when I know it will go further elsewhere.

I just wish they had that €100 subscription tier, for the equivalent of €1k pay-as-you-go use.

AntonJidkov 4 hours ago

Curious that it scores higher than opus 5.5 in cybersecurity because the closed models refuse to comply. I wonder if that means it's more susceptible to offensive uses.

redanddead 9 minutes ago

The closed providers are serving useless, bricked models that are tainted with their shitty system prompts

DevKoala 6 hours ago

One step closer to Le Chaton Fat.

PoignardAzur 6 hours ago

Can't wait until Mistral starts adopting fancy product line names.

Soon we'll have Mistral 6 Chaton, Mistral 6 Guépard, Mistral 6 Tigre, Mistral 6 Dents-de-sabre, Mistral 6 Beast King, etc.

scrollaway 4 hours ago

Panther, Jaguar … Wait, wrong company.

cedws 5 hours ago

Le Chaton Fat development was cancelled internally after it broke loose into the treat box.

pizlonator 7 hours ago

Refreshing to see this.

The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!

But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).

So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.

james2doyle 7 hours ago

You are quoting the discount pricing. It is 50% off for the next two weeks. The blog post has the real pricing up front in the card on the right side: https://mistral.ai/news/mistral-large-4/

After that, it will be much closer to GLM 5.3, but you can also get 5.3 in their API! I dont see people really talking about that.

pizlonator 7 hours ago

Good point.

drbscl 7 hours ago

I still need to evaluate it for my own workloads, but if you trust the benchmarks, seems about on par in quality vs GLM 5.3 Flash

Source https://artificialanalysis.ai/models/mistral-large-4?total-c...

barrell 7 hours ago

That is the promotional price, after which it doubles :(

pacha3000 5 hours ago

I've been dreaming of this for a simple reason: the french prose combined with GLM 5.3 reasoning capabilities.

GLM 5.3 is incredible because for the first time with an open-source model, it feels.. enough. I don't need much anymore, this model is great in everything. Except a thing : speaking french.

If the benchmarks are true, I'd be glad to switch entirely to Mistral.

rglover 7 hours ago

Excited to try this. The low costs v. benchmarks alone here are worth a serious test. K3 has been my daily driver for a month or two now and it's dramatically reduced token spend (while not having much of a negative impact on productivity).

This was the era of the AI race I was waiting for.

netvarun 7 hours ago

Curious now that SOL is cheaper than k3 - is k3 still your primary workhorse?

rglover 6 hours ago

Haven't tried it yet. It looks like sol is a hair more expensive on input, a hair less expensive on output ($2/in, $10/out per 1m, K3 = $0.82/in, $13/out per 1m).

bartstp 7 hours ago

How could I resist switching to a model named after my cat!?

laserbeam 7 hours ago

Stats be damned irrelevant. The naming is good with this one!

volkk 6 hours ago

you and 10k other redditors

wyrdcurt 3 hours ago

Maybe I'm missing something. Doesn't seem super impressive to me. A proprietary model with performance comparable to GPT-6 Luna and Deepseek 4.1 Flash, but at a higher price than either. The main selling point is that it's made in Europe... not very compelling, globally. I suppose maybe there is some niche where European-hosted open-weight models aren't enough to satisfy some EU regulation, where only the use of European-trained models is in compliance, but as a non-European I have no idea what that niche would be.

Side note: Wish this thread was more focused on talking about the model instead of debating about China and America. Whatever happened to staying on-topic?

ulimn 3 hours ago

You yourself also kind of pointed out why the discussion about USA and China is not off-topic. The niche Mistral wants to fill (afaik) is that it's European. I, as a European am really happy that they made such progress in so much worse (financial) conditions. I guess it's mostly geopolitics.

morningsam 13 minutes ago

It's proprietary only until the end of the month, when its weights will be released.

ktosobcy 5 hours ago

Awesome!

All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)

Narciss 8 hours ago

Le Chaton Fat is here!

HelloUsername 8 hours ago

Narciss 8 hours ago

Yeah just saw that, I'm gonna keep converting it in my head.

nsbk 8 hours ago

Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.

Tais-toi et prends mon argent!

quadruple 8 hours ago

I believe it is already available in the API no?

objektif 8 hours ago

Direct through them?

newtonsmethod 7 hours ago

angristan 7 hours ago

geroge_kyaw 3 hours ago

Still second most expensive open-weight model. I don't care about cybersecurity index. And still can't beat Chinese models but good to see European in the game.

xpct 7 hours ago

I don't know why, but I personally find Mistral's marketing strategy much more appealing than that of other companies.

For example, there's something about Anthropic's picked design and their little Claude avatars that's unsettling to me.

booty 6 hours ago

Sans context, I really like the Anthropic and Claude's faux-academic, minimal-ish, intellectual-ish branding.

(Especially the way it looked ~12 months ago -- it's gotten more cluttered since then. Perhaps unavoidably, as the breadth of their offerings has grown)

But over time it's begun to feel like unsettling cognitive dissonance as their ambitions grow and the stuff to worry about has piled up.

fodkodrasz 6 hours ago

> faux-academic, minimal-ish, intellectual-ish branding

By this you sourely don't mean the messages shown in Claude Code, where Pi would show "Working..."

Academic style: Cooking... Sautéeing... Julienning... and similar annoying faux-brogrammer moody status messages.

nine_k 5 hours ago

isoprophlex 6 hours ago

The logo isnt a stylized butthole. That sure helps endear me to them.

monocularvision 5 hours ago

I know the whole “AI logos look like buttholes” thing is a joke for most people, but I really speaks to the pornification of our society. I would never have thought “butthole” looking at any of their logos.

trollbridge 5 hours ago

b473a 5 hours ago

Y_Y 5 hours ago

smusamashah 6 hours ago

If it looks like that you, you will be seeing it everywhere.

troupo 6 hours ago

LarsDu88 5 hours ago

Why was I about to write this exact same comment, and why did you beat me to it?

financetechbro 5 hours ago

You are assuming that OP is not a fan of buttholes

fidotron 7 hours ago

The entire Anthropic branding is religious kitsch - deeply off putting, but apparently quite reflective of their reality.

walrus01 7 hours ago

I think there's such a thing as throwing too many marketing and sales people and too much polishing and "refinement" at something. I see in new product announcements from Microsoft as well. It's like seeing someone try too hard to impress you.

InsideOutSanta 6 hours ago

They've polished all personality out of their companies.

oytis 7 hours ago

And their cookie banner. Never thought I would like a cookie banner

YeahThisIsMe 6 hours ago

I do my best to avoid any AI marketing because I extremely despise it. I just haven't quite decided on my new profession, yet, but it's either going to be with plants or with animals.

fancyfredbot 7 hours ago

Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.

Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.

Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...

karannb 6 hours ago

From Guillaume

> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.

https://x.com/GuillaumeLample/status/2107461898127954001

Luker88 8 hours ago

Mistral Large 4: 1050B, 49 Active

GLM-5.3: 753B, 40 Active

I was hoping for something that hinted at smaller models too, but I guess not.

Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.

rahen 7 hours ago

Give them time. ML 4.0 was just pretrained. Mistral will certainly use it as the base for distillation and RL for smaller, better, more efficient iterations, just as the competition does.

gandreani 8 hours ago

It's competitive!

Good enough to show competence, and instill confidence in the team/company. Later releases can be more efficient.

I think it's a great release with that framing.

nolok 4 hours ago

It's Mistral Large, they usually publish Medium and Small later

baq 7 hours ago

just keep RL frying it should get better...

mchusma 7 hours ago

Something looks off in artificial analysis. Benchmarks aren’t everything, but not even close to the Pareto https://artificialanalysis.ai/models/mistral-large-4?cost=in...

I guess lots of token usage.

ianpurton 7 hours ago

Off Topic - The Mistral website - Really nice design. My guess, built by a human.

eternauta3k 6 hours ago

The key factor isn't whether the author used an LLM, but whether they had taste and attention to quality and iterated accordingly.

jvwww 5 hours ago

My guess would be designed by a skilled human with help from AI and built with AI by a skilled programmer.

1e1a 2 hours ago

Title is missing the official model name (Le chonk)

cbg0 8 hours ago

Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)

juliennakache 6 hours ago

Is that a reference to LeChuck in Monkey Island? Love that game!

square_usual 6 hours ago

No, it's a play on the Twitter/Reddit memeing about "Le Chaton Fat", see: https://www.reddit.com/r/MistralAI/comments/1u6f0dm/what_is_...

You'd be surprised how much of the AI world is fueled by memes.

cyanydeez 6 hours ago

litterally, that's what an LLM is doing?

jkingsman 6 hours ago

I think it's actually a wink and a nod towards the social media meme of "Le Chaton Fat," a fictional model that is jokingly attributed to Mistral, usually with century-defining benchmarks and unfathomable size.

https://x.com/i/trending/2066326562623127678

skc 8 hours ago

We're probably fast approaching the scenario where the cheapest models will win out.

aeneas_ory 8 hours ago

Benchmarks are better than expected! And probably got there without distillation ;)

water-drummer 8 hours ago

Is there a reason to believe why they wouldn't distill locally running open weights Chinese models?

mcbuilder 8 hours ago

Looks like they are doing 50% off to stay price competitive with DS Flash V4.1

segmondy 7 hours ago

I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!

dom96 8 hours ago

Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.

Why isn't this on openrouter yet? Is there a better router that gets models much quicker?

1 - https://bench.killswitch-lang.org/

dverlaeckt80 7 hours ago

Good to see Europe is at least a little bit still in the game.

https://docs.mistral.ai/inference/model-selection-guide?mode...

Cost is stated at half the price of GLM-5.3, which is quite interesting.

walrus01 8 hours ago

The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.

bloodmoon an hour ago

dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.

MrDrMcCoy an hour ago

Because screw Anthropic and OpenAI, that's why.

Kim_Bruning 7 hours ago

I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.

Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.

Not suitable for my purposes I don't think.

brendong 5 hours ago

Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?

andhuman 7 hours ago

At the end of the blog post we get this nugget.

> The pace of progress from here will be fast. Stay tuned.

scirob 4 hours ago

no hugging face link :( ... but hey its on openrouter yay

Here are my results

https://dach.peerbench.ai/compare?models=mistralai%2Fmistral...

Looks like a bit better than the recent Kolibri-1 but still below Qwen3.8 27B

XCSme 5 hours ago

I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...

ktosobcy 5 hours ago

Awesome!

All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)

aennassiri 7 hours ago

Excited to see this! Nice that they are saying this is just a first step.

Give them more compute!

valzam 7 hours ago

What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.

quadruple 7 hours ago

They explicitly lean in to cyber work, and they appear to be very permissive from their marketing:

> This is particularly important in cybersecurity, where provider-level refusals can block legitimate vulnerability research and incident response, and where losing access to a capability mid-incident can itself become a critical security risk. ML4 pairs top-tier cyber performance with open weights and self-deployment, giving organizations both the capability and the autonomy to run advanced security work under their own policies.

> That top score reflects a practical advantage. Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task. Yet defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block. This matters even more as threat actors increasingly jailbreak those same models to support offensive cyber activity

rarisma 8 hours ago

le chaton fat is real, my life is complete. Benches look crazy good for 1T.

whatever1 3 hours ago

I love the name! Teasing the ones making fun of them.

jrflo 7 hours ago

Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.

blauditore 5 hours ago

What's up with the name? It reminds me of my teenage self trying to speak in funny memes.

armaghanraza 4 hours ago

Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.

randomNumber7 4 hours ago

> Marketing decides which company makes it to the phones or PCs.

Right now a big part of LLM market is people using it for professional software development. Most of these users probably care about the quality of the model and also notice it during daily work.

For normal consumers, shure it doesn't matter. In the end the ai summary of google will probably be the most used as they are already exposed to it anyways.

carodgers 5 hours ago

Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.

Roark66 8 hours ago

This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.

I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.

I'd live to have one like that but EU made.

sixhobbits 8 hours ago

if it's not available yet why have a 'try it today' header at all?

> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.

newtonsmethod 8 hours ago

You can try it on their website, https://console.mistral.ai/playground

It's just that the open-weights aren't yet available (although the long delay is slightly annoying).

timcobb 7 hours ago

> Trained from scratch

How are they training without pirating the Z library corpus and all that?

sigmar 6 hours ago

I interpret "scratch" to mean brand new weights. Not that they aren't training on a corpus of human text

timcobb 6 hours ago

Right but how did they get a corpus, how do they compete legally without distillation?

sourcecodeplz 3 hours ago

at 200M tokens for the full AI suite run its not token efficient at all

AM1010101 4 hours ago

Half price on open router right now

ThouYS 8 hours ago

Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats

drbscl 7 hours ago

So about 2 or 3 generations behind, just like they were a year ago?

Nux 6 hours ago

Number one in Sovereign AI. Join our Discord.

tosh 8 hours ago

sorting the charts like that gives off weird vibes

https://mistral.ai/news/mistral-large-4/

ambicapter 8 hours ago

sorting the chart like what? You just linked to the main page.

ocamoss 7 hours ago

If you scroll down, most of the charts on that page are sorted s.t. Mistral's bar is right next to the worst competitor model, while the best competitor model's bar is positioned on the opposite side.

If one were to be cynical one could say that it's intentionally making Mistral's result look better than it actually is by making it harder to compare the bar heights.

EDM115 7 hours ago

We actually got Le Chaton Fat before GTA 6

tdubey 8 hours ago

Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?

lucrbvi 8 hours ago

Space Bunny Alpha is probably MiniMax M3.1 (rumors on Twitter since it seems to have a similar tokenizer).

irl_zebra 8 hours ago

Yes broad consensus had developed in the ten minutes between announcement and you asking if consensus had developed, and I'm excited to report that it consensed in the affirmative -- it IS Space Bunny Alpha!

tosh 8 hours ago

sorting the charts like that gives off weird vibes

jasonjmcghee 8 hours ago

Had the same thought - feels chart crime adjacent

lern_too_spel 7 hours ago

Bar charts should start at zero. If they don't start at zero, there should be a clear visual indicator that the chart has been trimmed without having to read the axis labels. I hate that this has to be repeated so often that it has become a cliché.

chriskanan 4 hours ago

I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.

Mistral doesn't publish the science.

armaghanraza 4 hours ago

Brother I think No one publishes the science

igleria 7 hours ago

I thought lechonk motto was just a meme!

maz1b 7 hours ago

I'm glad they're keeping at it!

4rtem 8 hours ago

Previous one is barely in top 50 on arena.ai

theturtletalks 8 hours ago

A bit disappointing to see it still lagging behind Chinese open models. Those Chinese models are pushing proprietary models to raise the bar, but we need equally strong non-Chinese open models to challenge the Chinese ones in turn.

pietz 8 hours ago

I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.

ianpurton 8 hours ago

On Prem. Thats a bid deal for some enterprises.

also the benchmarks are not necessarily indicative of how well the model will perform in its own harness with its own skills.

pietz 7 hours ago

Any open weights model is "on prem".

jvwww 5 hours ago

I mean usually the benchmarks make any model feel better than how they actually perform.

redanddead 7 hours ago

better than K3 and DS4, cool

theanimeshs 6 hours ago

impressive release this time by Mistral. bullish.

gabe-santana 6 hours ago

Amazing! just tested

staticman2 8 hours ago

Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.

throwa356262 8 hours ago

It certainly has helped OpenAI and Anthropic get their KV cache costs under control.

Tade0 8 hours ago

It's no secret that everyone is dis-stealing from everyone else.

drbscl 7 hours ago

I don't see how distillation relates to using the published techniques developed by Deepseek, Moonshot, Zhipu, etc

keithnoizu 8 hours ago

touche

TokiBot 3 hours ago

Where can it be tested?

maxdo 8 hours ago

Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well

LoganDark 4 hours ago

1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!

saberience 8 hours ago

Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.

https://www.vals.ai/benchmarks/vals_index

ofirg 8 hours ago

where does sit on the pareto distribution compered to Le Chaton Fat?

glerk 4 hours ago

Massive fumble not to call it “le chaton fat”.

0xbadcafebee 5 hours ago

Too bad this got marked as a dupe, as it actually has benchmark info unlike the other page which is just docs.

The weird thing is how worse they are at things like coding than other open weights. You'd expect them to at least distill coding from other open weights to match them.

alpineman 8 hours ago

>> Unofficially ML4, very officially: le Chonk

Honestly just nice to see a leader in this space not take themselves so seriously.

jvwww 5 hours ago

Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.

htrp 7 hours ago

europe finally getting into the race here.

Razengan 7 hours ago

Awh I was half expecting a zombie pirate..

scrubmunch 8 hours ago

wowza le models a heckin chonker

clavicle1009 8 hours ago

YES finally

petesergeant 8 hours ago

Anyone have any indication when I can get my hands on a developer plan for this?

Havoc 6 hours ago

They do sell subscriptions I think

spwa4 8 hours ago

Don't believe Mistral. They're wrong. It's really called "Le chaton fat".

Also 1T-A49B. Weights currently closed but promise to open source them by the end of the month.

Great release movie.

alterom 8 hours ago

>Don't believe Mistral. They're wrong. It's really called "Le chaton fat".

OpenAI's therapist: Le Chaton Fat isn't real and cannot hurt you

Le Chaton Fat:

crimsoneer 8 hours ago

Woah, this seems like a big deal (assuming the benchmarks are as good as claimed)?

Mistral slightly proving me wrong (and I'm not mad).

baggachipz 8 hours ago

Now THAT'S how you name a model. Take note, others.

charcircuit 5 hours ago

>Frontier performance

* proceeds to not compare to Opus 5.5

super256 5 hours ago

They say "open weights frontier performance". Of course they aren't comparing to Anthropic, because they are not putting their weights online.

charcircuit 3 hours ago

I'm talking about the section of the article labeled "Frontier Performance", which does not specify that.

retinaros 7 hours ago

quick question why put GLM 5.3 at 61 while a quick check on DeepSWE 1.1 puts it at 69?

also they forgot muse spark at 75% while claiming they were outshining all US models?

bdcravens 8 hours ago

Can we consolidate the posts? Currently there's 3 on the front page, basically all pointing to Mistral's messaging in different places.

thomastraum 4 hours ago

Mistral wont win the AI race because of the model names. I wont bother an arrogant Parisian hipster with my insecure prompts who then plays with his moustache and responds with a judgmental "pfff"

sinan-faizal 7 hours ago

i subbmitted a partnership proposal in your contact.

erichocean 8 hours ago

Does Mistral ever advance the state of the art on any dimension?

And if not, why do they exist?

Update: The number of people advocating not innovating is wild. There is no reason why Mistral cannot innovate in ML, they explicitly choose not to. My point is that, given that choice, they should spend their GPU hours differently.

"Sovereign AI" is a joke, there is no substantive difference between a post-trained open weight model from an American or Chinese company and what Mistral is doing today, beyond spending 80% of their GPU hours reproducing a last-gen model's pretraining.

CJefferson 8 hours ago

Why would the Europeans want a high-quality open source model that isn't either owned by American Trillion-dollar companies or created by the Chinese? You really can't think of any reasons?

erichocean 8 hours ago

> high-quality open source model

> that isn't owned

I don't know what definition of open weights you are using, but the "open" part is what allows you to use it without being beholden to American Trillion-dollar companies or the Chinese.

Mistral could do exactly what Cursor does, and post-train the model to meet the needs of Europeans.

That's a far more efficient (and useful!) use of GPUs than yet another pretrain on an outdated model backbone.

eigenspace 8 hours ago

Has the French military advanced the state of the art on any dimension in the last hundred years?

And if not, why do they exist?

pizzafeelsright 8 hours ago

Second hand rifle distributor?

azan_ 8 hours ago

Of course they have!

eigenspace 8 hours ago

swiftcoder 8 hours ago

So that there exists an EU-native option in the near-frontier LLM space?

Not everyone is wild about being downstream of either the Chinese or US governments, particularly when it comes to things like cybersecurity

embedding-shape 8 hours ago

Just like Microsoft and most companies around PCs, their existence isn't about pushing any boundaries, but seemingly targeting deploying what's already been "invented" into the corporate world and bigger companies. Not "wrong", just different.

erichocean 8 hours ago

Sure, agreed. But they do their own pre-training, at great expense, on outdated model backbones.

Wouldn't it be better to do something more like Cursor, and RL on an existing pretrained model if you're not innovating anyway?

notfromhere 7 hours ago

esafak 8 hours ago

Mistral is one of the few European AI labs. Look up "sovereign AI".

whstl 8 hours ago

Does erichocean ever advance the state of the art on any dimension?

And if not, why do they exist?

erichocean 8 hours ago

Yes, actually. Thanks for asking.

goatley 7 hours ago

Finally, a model small enough to self-host on my 2012 MacBook Air if I don't mind my house reaching room temperature in 2 seconds.

alex_duf 7 hours ago

You might be missing three 0s on the parameter count, or what am I missing?

I'm assuming you need somewhere 0.5 to 1TB of RAM for the weights only

goatley 6 hours ago

not to ignore you but how is it possible that I have zero karma

randomNumber7 4 hours ago

londons_explore 5 hours ago

Still a long way behind closed models sadly :-( The gap between open and closed seems the biggest it's been for a year or two.

whatsThisBtn4 2 hours ago

I try to stay up on AI models, but I've given up on Mistral. I have tried using it too many times and it's the worst out of major models. Even Kimi and DeepSeek are miles ahead.

Seems like it's another European company that is only alive through government financing.

I understand wanting a home grown industry, but with AI models, just abliterate a DeepSeek model.

Total misunderstanding of the industry. Europe needs hardware, not a model that needs to be replaced monthly.

bigyabai 2 hours ago

> Total misunderstanding of the industry. Europe needs hardware

With all due respect, I have no idea what you are trying to say with this comment and have no clue how "hardware" would turn Europe's tides. You explained nothing.

Do they need training hardware? Inference hardware? ASICs or GPGPUs? Edge hardware? Agent hosts? Faster cores, or wider ones? Taller systems, or more parallel ones?

Your vagueness completely undermines the authority that your criticism relies on.