DeepSeek V4 Flash on a Single AMD MI300X (github.com)

325 points by zhoutong 10 hours ago

majke 10 hours ago

I don't think you can buy a single "MI300X" unit, right? Only the box with x8 of these at a cost of ~250K EUR.

zhoutong 9 hours ago

It’s available on demand from a few cloud providers. Seems like the cheapest is AMD Developer Cloud (https://www.amd.com/en/developer/resources/cloud-access/amd-...) powered by Digital Ocean at $1.99/hour.

Edit: Now I think about it, this might be the cheapest way to run the DeepSeek V4 Flash 0731 on a dedicated inference server at original weights. I haven’t run mixed load benchmarks but I guess it’s possible to generate $3-$4 worth of tokens per hour and still maintain a usable per-user throughput.

WASDx 9 hours ago

At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.

zhoutong 5 hours ago

lnenad 8 hours ago

Almondsetat 8 hours ago

krisknez 8 hours ago

thrownaway561 8 hours ago

Tepix 7 hours ago

If you have 2x DGX Spark it will run quite nicely. They cost only $8000 or so and use less power so you may be able to rent them cheaper than the MI300X.

I found an offer to rent two at $1.65 per hour https://spark.enverge.ai/#pricing

The MI300X will vastly outperform it for only a slightly higher price.

K0IN 4 hours ago

langs 9 hours ago

You need to optimize the KVCache part(save to disk to save compute) to achieve this goal.

throwawayffffas 6 hours ago

You can get one on ebay for like 20k, but it comes without the backplane and i dont think there is a pcie card adaptor from china like the ones for h200.

Lwerewolf 9 hours ago

The MI350p exists and should run a decent quant (say, the ~96GB antirez mix) well, but you can get two rtx pro 6000s for one of these, or 8x (actually more) r9700 + probably the gear to run them, etc.

Otherwise, you can probably buy one of these second hand from somewhere (SXM A100s are available that way) and run it in an adapter board.

touisteur 6 hours ago

I thought MI350P wasn't available yet, curious where to source it right now.

FuriouslyAdrift 2 hours ago

_joel 8 hours ago

I thought it was a consumer grade GPU until I saw the 192GB of HBM and 256GB or RAM.

FuriouslyAdrift 2 hours ago

It's a chopped down MI350X (roughly half the performance)

varispeed 8 hours ago

To be fair the development of GPUs have stalled over the years. If they kept up with the progress instead of focusing on enterprise market, likely 256GB consumer GPU would be a norm today.

baalimago 9 hours ago

Give it an AI-bubble pop and these will be flooding the market.

segmondy 7 hours ago

no they won't , the bubble is a financial thing. the demand is real and not going away.

atwrk 5 hours ago

_factor 9 hours ago

They will be instantly bought out by companies, not individuals. The consumer bubble won’t pop for quite a while yet. Production also won’t ramp up while lack of real competition keeps the demand high.

aurareturn 9 hours ago

When is it popping? Is the AI bubble in the room with us now?

dghlsakjg 6 hours ago

amrit3128 8 hours ago

baalimago 8 hours ago

slaw 8 hours ago

tamimio 7 hours ago

Thing is, GPUs will always be on demand, look at their history, initially for gaming, then for hash cracking, then 3D rendering, then for crypto mining, and now AI training and fine tuning. When AI bubble bursts, there will be another bubble taking over.

The only solution is more companies making high end units, only competition will make it better for consumers.

GTP 7 hours ago

Strange that in the prior art they didn't list DwarfStar, as it is able to run the same model (probably quantized differently though) in less memory. Maybe the author isn't aware of it?

MaKey 5 hours ago

The focus of this repo is the MI300X and DwarfStar doesn't include any optimizations/fixes for it.

wmf 2 hours ago

https://github.com/antirez/ds4/pull/484

I assume this is parallel work.

fergusfinn 6 hours ago

nice! i think the higher HBM on Mi300x is really useful for this kind of thing

we did some work on this for 2xMi300x (kindly referenced in the readme) https://blog.doubleword.ai/deepseek-v4-flash-mi300x. https://hotaisle.xyz/quick-start hotaisle is great for getting Mi300x to experiment with

Tepix 7 hours ago

Unfortunately, the MI300X is an OAM module. The MI350P is the one you want: It's a PCIe card, but it has less memory: 144GB.

Luckily, DeepSeek V4 Flash will run in 144GB too because it's 256 MoE exports are native MXFP4 quantized.

WhitneyLand 6 hours ago

How do you figure that?

When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.

And this implementation is already cutting down the 1M token context window you would normally get.

Tepix 42 minutes ago

For sure if you want to properly utilize the model with several users in parallel and large context you'll want two MI350P.

craftkiller 3 hours ago

Just want to add that while the MI350P is a PCIe card, it is designed for servers. It just has a heatsink (with no fan) which the powerful full-case fans of a rackmount server are supposed to cool. So while the MI350P is certainly more attainable for us regular folk due to its formfactor, we won't be able to just drop it into our gaming PCs like a regular graphics card.

That being said, if you're dropping tens of thousands of dollars on graphics cards then picking up a rackmount case to go around the card is pretty insignificant.

cyberax 2 hours ago

Just put a fan on it. It's just 600W, so nothing super-special is needed. Or add a water cooler.

monster_truck 2 hours ago

WhitneyLand 6 hours ago

Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”.

Dumbed down quantization?

No. Full intended inference weights preserved, so far so good.

Slow performance?

No again. Looks like you could get over 150 tokens/second.

Give up context window size?

Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.

monster_truck 2 hours ago

In my experience the 1M context is genuinely too much. The first time I swapped from OAI to DSv4P, I checked and double checked that the harness/etc was working correctly over the course of hours and hours of work thinking that I had set something up wrong because it simply never had to compact! The drop in quality is arguably less than that of what you get from compact to compact on Codex, which is good for what it is or was.

Was also surprised to learn just how much of Codex's window was being burnt on shit I didn't want or use. Sure I can pass this and that flag to eliminate most of it, but for a $200/mo product aimed at professionals, that isn't something anyone should have to janitor (also totally ignoring the bandaid of banked resets they've slapped over their repeated mistakes).

It's wild just how far $20 will get you with Deepseek, even at their new rates. Buyers Remorse is my very least favorite feeling, I felt sick thinking about what the $1200 I had given OAI this year would have gotten me had I only tried sooner.

bwfan123 5 hours ago

I am curious if there has been work to remove experts from an open-weights model. The goal would be to reduce the size to be able to run on desktop GPUs without compromising quality. For a focused usecase - say coding, you dont need a model that knows world history. And, I am not talking about quantization. If it is possible to determine which experts are active for some usecases, and surgically remove the others.

smallerize 3 hours ago

Experts aren't trained on separate tasks. More recent routers are designed to spread out requests even more evenly, and they were already pretty even.

Tepix 41 minutes ago

Yes. It's called REAP and from what I've seen, results aren't stellar.

monster_truck 2 hours ago

Glazing over a lot, that's how they work already, just not in the way you think. A relatively small fraction of the model is active at any given time

xorfish 7 hours ago

This is still quite a bit away from the performance that deepseek gets on their H800. In their DSpark paper they report a throughput of 15k tokens/s/gpu. The MI300 should be able to compete with the H800 so there are probably still quite a few optimizations that can be made.

somnial 4 hours ago

throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare.

also it would be interesting to figure from the DSpark paper whether their numbers are consistent with the GPUs still being H800s, since they never actually say...

PrimeAli 2 hours ago

Great

sylware 7 hours ago

Is their hardware programming interface reasonable for implementing inference of frontier models: no quantization, several tera params?

BTW, how many many params open weight frontier models have? A few teras, 100s of teras?

wmf 2 hours ago

Yes, ROCm can be used to run frontier models and is being used by OpenAI, Anthropic, and Meta.

sylware 25 minutes ago

I would prefer direct hardware kernel interface.

Like linux DMABUFs with userland hardware command ring buffers (I guess this hardware ring buffer instance would be specific to a VMID and a PASID).

wren6991 6 hours ago

Kimi-K3: 2.8T

Qwen3.8-Max: 2.4T

DeepSeek V4 Pro: 1.6T

DeepSeek V4 Flash: 284B

(all are total parameter counts, not active parameters)

sylware 37 minutes ago

Rumors say chatgpt/claude/gemini/etc are in the 100s of teras. True?

wren6991 25 minutes ago