DeepSeek V4 Flash on a Single AMD MI300X (github.com)
325 points by zhoutong 10 hours ago
majke 10 hours ago
I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.
zhoutong 9 hours ago
It’s available on demand from a few cloud providers. Seems like the cheapest is AMD Developer Cloud (https://www.amd.com/en/developer/resources/cloud-access/amd-...) powered by Digital Ocean at $1.99/hour.
Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.
WASDx 9 hours ago
At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.
zhoutong 5 hours ago
lnenad 8 hours ago
Almondsetat 8 hours ago
krisknez 8 hours ago
thrownaway561 8 hours ago
Tepix 7 hours ago
If you have 2x DGX Spark it will run quite nicely. They cost only $8000 or so and use less power so you may be able to rent them cheaper than the MI300X.
I found an offer to rent two at $1.65 per hour https://spark.enverge.ai/#pricing
The MI300X will vastly outperform it for only a slightly higher price.
K0IN 4 hours ago
langs 9 hours ago
You need to optimize the KVCache part(save to disk to save compute) to achieve this goal.
throwawayffffas 6 hours ago
You can get one on ebay for like 20k, but it comes without the backplane and i dont think there is a pcie card adaptor from china like the ones for h200.
Lwerewolf 9 hours ago
The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.
Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.
touisteur 6 hours ago
I thought MI350P wasn't available yet, curious where to source it right now.
FuriouslyAdrift 2 hours ago
_joel 8 hours ago
I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.
FuriouslyAdrift 2 hours ago
It's a chopped down MI350X (roughly half the performance)
varispeed 8 hours ago
To be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.
baalimago 9 hours ago
Give it an AI-bubble pop and these will be flooding the market.
segmondy 7 hours ago
no they won't , the bubble is a financial thing. the demand is real and not going away.
atwrk 5 hours ago
_factor 9 hours ago
They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.
aurareturn 9 hours ago
When is it popping? Is the AI bubble in the room with us now?
dghlsakjg 6 hours ago
amrit3128 8 hours ago
baalimago 8 hours ago
slaw 8 hours ago
tamimio 7 hours ago
Thing is, GPUs will always be on demand, look at their history, initially for gaming, then for hash cracking, then 3D rendering, then for crypto mining, and now AI training and fine tuning. When AI bubble bursts, there will be another bubble taking over.
The only solution is more companies making high end units, only competition will make it better for consumers.
GTP 7 hours ago
Strange that in the prior art they didn't list DwarfStar, as it is able to run the same model (probably quantized differently though) in less memory. Maybe the author isn't aware of it?
MaKey 5 hours ago
The focus of this repo is the MI300X and DwarfStar doesn't include any optimizations/fixes for it.
wmf 2 hours ago
https://github.com/antirez/ds4/pull/484
I assume this is parallel work.
fergusfinn 6 hours ago
nice! i think the higher HBM on Mi300x is really useful for this kind of thing
we did some work on this for 2xMi300x (kindly referenced in the readme) https://blog.doubleword.ai/deepseek-v4-flash-mi300x. https://hotaisle.xyz/quick-start hotaisle is great for getting Mi300x to experiment with
Tepix 7 hours ago
Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB.
Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.
WhitneyLand 6 hours ago
How do you figure that?
When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.
And this implementation is already cutting down the 1M token context window you would normally get.
Tepix 42 minutes ago
For sure if you want to properly utilize the model with several users in parallel and large context you'll want two MI350P.
craftkiller 3 hours ago
Just want to add that while the MI350P is a PCIe card, it is designed for servers. It just has a heatsink (with no fan) which the powerful full-case fans of a rackmount server are supposed to cool. So while the MI350P is certainly more attainable for us regular folk due to its formfactor, we won't be able to just drop it into our gaming PCs like a regular graphics card.
That being said, if you're dropping tens of thousands of dollars on graphics cards then picking up a rackmount case to go around the card is pretty insignificant.
cyberax 2 hours ago
Just put a fan on it. It's just 600W, so nothing super-special is needed. Or add a water cooler.
monster_truck 2 hours ago
WhitneyLand 6 hours ago
Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.
Dumbed down quantization?
No. Full intended inference weights preserved, so far so good.
Slow performance?
No again. Looks like you could get over 150 tokens/second.
Give up context window size?
Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.
monster_truck 2 hours ago
In my experience the 1M context is genuinely too much. The first time I swapped from OAI to DSv4P, I checked and double checked that the harness/etc was working correctly over the course of hours and hours of work thinking that I had set something up wrong because it simply never had to compact! The drop in quality is arguably less than that of what you get from compact to compact on Codex, which is good for what it is or was.
Was also surprised to learn just how much of Codex's window was being burnt on shit I didn't want or use. Sure I can pass this and that flag to eliminate most of it, but for a $200/mo product aimed at professionals, that isn't something anyone should have to janitor (also totally ignoring the bandaid of banked resets they've slapped over their repeated mistakes).
It's wild just how far $20 will get you with Deepseek, even at their new rates. Buyers Remorse is my very least favorite feeling, I felt sick thinking about what the $1200 I had given OAI this year would have gotten me had I only tried sooner.
bwfan123 5 hours ago
I am curious if there has been work to remove experts from an open-weights model. The goal would be to reduce the size to be able to run on desktop GPUs without compromising quality. For a focused usecase - say coding, you dont need a model that knows world history. And, I am not talking about quantization. If it is possible to determine which experts are active for some usecases, and surgically remove the others.
smallerize 3 hours ago
Experts aren't trained on separate tasks. More recent routers are designed to spread out requests even more evenly, and they were already pretty even.
Tepix 41 minutes ago
Yes. It's called REAP and from what I've seen, results aren't stellar.
monster_truck 2 hours ago
Glazing over a lot, that's how they work already, just not in the way you think. A relatively small fraction of the model is active at any given time
xorfish 7 hours ago
This is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.
somnial 4 hours ago
throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare.
also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs still being H800s, since they never actually say...
PrimeAli 2 hours ago
Great
sylware 7 hours ago
Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?
BTW, how many many params open weight frontier models have? A few teras, 100s of teras?
wmf 2 hours ago
Yes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.
sylware 25 minutes ago
I would prefer direct hardware kernel interface.
Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).
wren6991 6 hours ago
Kimi-K3: 2.8T
Qwen3.8-Max: 2.4T
DeepSeek V4 Pro: 1.6T
DeepSeek V4 Flash: 284B
(all are total parameter counts, not active parameters)
sylware 37 minutes ago
Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?