Xiaomi MiMo v2.6 (mimo.xiaomi.com)
278 points by volf_ 2 hours ago
rao-v 2 hours ago
I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
MangoCoffee 23 minutes ago
maybe this is why Dario want to slow down AI development and all the big AI labs in the USA is singing the same song.
whey they all singing the same tune. it make me question what is their real motives.
they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.
jwolfe 5 minutes ago
Please explain how putting an upper bound on how good the strongest models can be prevents cheaper less strong models from catching up, rather than enabling it. I do not understand this argument at all.
rbjorklin a few seconds ago
bellowsgulch a minute ago
earthnail an hour ago
Thanks so much for sharing this. As someone who mostly watches from the sideline, can you share what you can see in this dashboard that someone like me can't see? Is it the metrics themselves that they measure (the metrics tab is absurdly detailed), something in the notices, or something else I missed?
rao-v 23 minutes ago
I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.
Here is one really neat bit:
A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".
Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.
(if you've ever tried to teach humans new skills, you’ve probably noticed this too!)
When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).
Lo and behold, the dashboard literally has:
partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)
train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)
How cool is that?!
And don't get me started on the clever ideas hiding behind dynsam/avg@n ...
dgellow a few seconds ago
tancop 33 minutes ago
The best thing they did is being open about all the setbacks they had to deal with. They logged every restart with a reason, talked about dropping a cyber dataset after it degraded coding benchmarks. Also published real time training loss, benchmark scores after every checkpoint and running cost estimates.
Really the only thing missing was dataset descriptions, the dashboard only had random IDs like "dataset-zrso". I guess it's their lawyers fault.
verdverm an hour ago
the existence, who else has a live dashboard for the RL late-training?
ignoramous 16 minutes ago
> got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores
Xiaomi MiMo is led by Luo Fuli, a former Alibaba & DeepSeek employee. Perhaps it is due to Luo just how similar Xiaomi's tech & GTM approach is to DeepSeek's.
- How Luo Fuli Keeps an Earthy Touch as she Soars Through the AI World, https://newsen.pku.edu.cn/news_events/news/people/15385.html (https://archive.vn/I8Pmu).
- Luo Fuli, the 30-year-old ‘AI genius girl’ behind DeepSeek’s success?, https://e.vnexpress.net/news/tech/personalities/who-is-luo-f... (https://archive.vn/sb3B6).
lwansbrough an hour ago
Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.
tacomagick an hour ago
Absolutely! Chinese models are both cheaper and more capable in many cases, compared to the American models and their makers continuously fumbling or reducing model capability with each update. Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.
user43928 an hour ago
OpenAI decreased prices with the 5.6 model family.
And later they further cut Sol and Terra pricing by 20% (maybe only in the API) and Luna by 80%.
In fact Luna still outperformed DeepSeek Flash 4.1 in cost per task on Artificial Analysis when I last checked.
However, Luna is slightly less intelligent. I have a feeling that it's pretty dumb and prone to hallucination unless running at xhigh or max effort, where it somehow manages to work quite well.
I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.
The competition is great, and I hope Chinese models will continue to force leading US labs to offer models at a low price point.
That said, I don't think the Chinese labs have anything over OpenAI and Anthropic when it comes to capability or efficiency - I have no reason not to believe the US labs have even lower cost to serve the models.
Implicated 4 minutes ago
tacomagick 42 minutes ago
goosejuice 34 minutes ago
> Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.
OpenAI reduced prices and Anthropic increased weekly usage limits.
SyneRyder 44 minutes ago
Yep, I'm trending in that direction, and I'm someone with Claude stickers all over my laptop. My main app dev work is still going to Claude, but everything else is going to China even at API rates now.
One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.
rapind 11 minutes ago
API rates still aren’t quite competitive with the OpenAI x20 accounts, but they are definitely getting close with deepseek 4.1 flash. I spent a few days with only 4.1 and was very impressed.
verdverm an hour ago
I have a contrarian opinion that China passing America in Ai is the Sputnik moment we need to leave the hubris behind and get our mojo back
debatable if a turn around is possible before '29
swingandamiss an hour ago
No, because I'd rather not support our economic and military rivals.
lwansbrough an hour ago
I'm Canadian so this sentiment has little value in 2026 unfortunately.
zemvpferreira 2 minutes ago
ActionHank an hour ago
tancop 24 minutes ago
rayiner 36 minutes ago
Freedom2 an hour ago
Agreed, and also because I support freedom of speech!
girvo an hour ago
simonw an hour ago
Pelicans for Flash: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Pelicans for Pro: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
phainopepla2 an hour ago
I think we can say pretty confidently they aren't pelican-bench-maxxing
brcmthrowaway an hour ago
Just me, or do these look bad?
Qwen3.8-27b pelican was amazing on Mac.
idiotsecant an hour ago
Looking terrible isn't nessesarily a bad thing. The pelican is heavily pre trained now. Having a crappy pelican means you didn't try to juke the stats.
broodbucket 28 minutes ago
handfuloflight an hour ago
How does this translate to coding performance, which is what most of HN cares about (...I assume)?
simonw 43 minutes ago
It means they're good at writing SVGs, in particular SVGs of animals riding modes of transport!
lanyard-textile 39 minutes ago
I only visit HN for the pelicans, personally.
lukewrites 28 minutes ago
Imanari an hour ago
ish… at least we can be sure they don’t benchmaxx the pelicans lol
stymaar 2 hours ago
Flash[1]: 309B total / 15B activated parameters
Pro [2]:, 1.02T total / 42B activated parameters
verdverm an hour ago
There's also a Qwen 3.5 9B distill
gandreani an hour ago
Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
mydreamof an hour ago
verdverm an hour ago
curious why the HF pill (on the right) always has inaccurate values
stymaar an hour ago
I noticed the same, and I wonder as well.
verdverm an hour ago
segmondy 42 minutes ago
more like 500B in FP8
nemothekid 2 hours ago
Looking at the frontend design examples; why do these models seem to love the "01 - UPPERCASE TEXT" motif. It's everywhere now (see https://try.cloudflare.com/, which has '01 · QUICK TUNNELS', but no "02" anywhere).
danvayn an hour ago
My guess is that by function they break down frontend sections or components into pieces and I believe document things for themselves on some level, or purposely are verbose in this way. It is probably also shaped by users and existing web patterns. They probably get reinforced by models the more common they become.
pphysch 36 minutes ago
The extraneous small-caps labels are one of the main idiosyncrasies of AI generated markup. I wonder how much of this is a "scaffolding" technique to help the model build stable designs. But was it reinforced in RLHF or an emergent behavior of the models?
sandblast an hour ago
Nice catch!
user43928 an hour ago
I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
GPT 6 Astra 59.6
Claude Fable 5.1 55.1
Claude Opus 5 49.0
MiMo-V2.6-Pro 34.9
MiMo-V2.6-Flash 28.8
DeepSeek V4.1 Flash 26.8
MiMo-V2.5-Pro 1.5
ExploitGym GPT 6 Astra 42.4
Claude Fable 5.1 30.4
Claude Opus 5 22.1
MiMo-V2.6-Pro 17.8
MiMo-V2.6-Flash 6.0
MiMo-V2.5-Pro 0.1
DeepSWE v1.1 DeepSeek V4.1 Flash 74.2
Claude Opus 5 74.0
GPT 6 Astra 74.0
MiMo-V2.6-Pro 71.9
Claude Fable 5 70.0
MiMo-V2.6-Flash 67.9
MiMo-V2.5-Pro 19.0mokre an hour ago
Maybe you should not trust any of the benchmarks!
varispeed 43 minutes ago
They match my experience. Astra and Fable I rate below Sonnet. They are incredibly poor. They were excellent for a couple of days after release and then plummeted.
Maybe I am being routed to more quantised versions or less capable models with system prompt to fake Astra or Fable.
vatsachak 2 hours ago
Wow, the chinese labs are getting good at advertising model releases. The moat is thin.
Some features of the release I like:
- Demonstration of diverse tasks, such as using a DAW
- Graphs from various benchmarks and price ranges
- Real world use of the model in scientific environments
syntaxing an hour ago
All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
verdverm an hour ago
trvz 44 minutes ago
That’s for toy GPUs, like the 5090.
verdverm 39 minutes ago
brcmthrowaway an hour ago
Is there a gamechanger around the corner to reduce DRAM requirements?
zozbot234 an hour ago
You could always stream from SSD storage. Especially effective if you get a cheap old-gen HEDT with lots of PCIe slots to add NVMe storage to and reasonable overall PCIe bandwidth.
jkingsman 27 minutes ago
stymaar an hour ago
n-gram per-layer embeddings[1][2] might be it.
[1] https://sebastianraschka.com/llm-architecture-gallery/per-la...
[2]: See DS 4.1-Flash and Qwen-3.8-Next.
verdverm an hour ago
volf_ 44 minutes ago
I've got a working recipe to run this model on Dual DGX Spark: https://github.com/volfco/spark-vllm-docker/blob/main/recipe...
Averages ~25-35tok/s which isn't bad for a first attempt.
eriquesito 37 minutes ago
Funny that all but one video has audio, the house 3D model one, where you can hear (what I assume are) Xiaomi's engineers talking about who knows what.
ddxv 2 hours ago
This looks great in terms of cost and capabilities, truly pushing the frontier forward in terms of open weight light weight models.
thrownawaysz an hour ago
>Night 0.8x Usage, 00:00-08:00 -UTC+8
It's because offpeak electricity is cheaper?
Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm
MisterMunchkin an hour ago
I really liked MiMo 2.5, it was really affordable and actually had vision, unlike DeepSeek. (DeepSeek has only recently added it)
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
perrygeo 13 minutes ago
Can we afford to look past it? If/when claudeslop starts infecting every new model to such an extent, that model will produce its own slop, infecting new models... At what point do we lose all reliable methods for establishing "truth"? This is epistemic collapse waiting to happen. I honestly thought it would take longer... holding out for a coherent shared reality in 2030 seems optimistic.
omani an hour ago
how do you recognize "claudeslop"?
Bluestein an hour ago
It's an honest, load-bearing, simple thing.-
DanMcInerney 2 hours ago
This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
algoth1 2 hours ago
Finally a lab that doesn't cheat on the charts
bertili an hour ago
They mixed up DeepSeek 4.1 Flash with something else on this page, possibly DeepSeek 4.1 Flash means Gemini 3.8 Flash.
varispeed an hour ago
These benchmark are useless as they don't say whether they were done before or after Fable and Astra got nerfed.
gigatexal an hour ago
Leaning into what it cost to train is hilarious and an obvious shot at US frontier labs spending tens to hundreds of millions or more to train their models.
alfalfasprout an hour ago
The moat for OAI and anthropic seems to be very quickly shrinking. Chinese labs are now using RSI-like approaches and even without resorting to heavy distillation they're catching up in a couple of months vs. what would have been 6-12 months a year prior.
And as these models get better the pace of training is quickly speeding up too.
This doesn't bode particularly well for anthropic/OAI after they go public.
verdverm 43 minutes ago
token vendors are headed to the same place mobile data vendors went, this is good for everyone but those who thought they could maintain exorbitant prices
NooneAtAll3 an hour ago
does anyone know what unnamed model is on paretto frontier picture right between MiMo 2.5 and 2.6?
so weird to acknowledge someone being on the front edge, but not name it
AnodicElegy an hour ago
Pretty sure that's Luna xhigh.
spwa4 an hour ago
As for the stats that everyone wants:
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Perhaps with IQ2 flash will run on 128G M5?
omani 2 hours ago
ah, would you look at that. I was wondering why mimo 2.5 became "dumber" the last weeks. I was speculating they are probably about to release a new version of the model. because the model really acted out a lot. especially the last two weeks. dont know, was just a feeling, highly speculative.
but now I got my "proof".
sandblast an hour ago
I guess that would only be possible if your provider was Xiaomi itself?
omani an hour ago
yes. I use opencode and opencode uses Xiaomi as a provider.
jwpapi an hour ago
In the chart they use "Pareto Line", which I think is wrong. Pareto is 20% effort leading to 80% results. Which could be interpreted as models costing 20% having 80% of peak intelligence, but that’s not what it looks like to me.
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.
abound an hour ago
There are two (or more) concepts named after the same person:
- Pareto efficiency/Pareto curves: Basically the convex hull of points along the edge of a graph, indicating the best tradeoff between the axes. This is what the post is talking about.
- Pareto principle: this is the 80/20 rule you're talking about
nextaccountic an hour ago
No, Pareto refers to Pareto efficiency https://en.wikipedia.org/wiki/Pareto_efficiency
What you call "frontier line" is also called "Pareto frontier" https://en.wikipedia.org/wiki/Pareto_front
Your description of it is basically correct though
shmolyneaux an hour ago
This is the Pareto Front [1], rather than the Pareto principle. It's the idea that anything that's more intelligent is more expensive and anything that's less expensive is less intelligent.
hashmush an hour ago
"Pareto" is many things, but here it does indeed refer to the frontier: https://en.wikipedia.org/wiki/Pareto_front
jwpapi an hour ago
Thank you guys. I learned something new.