CUDA for AMD on Windows (github.com)
145 points by chiassedu80 14 hours ago
linuxhansl 11 hours ago
Off-topic and somewhat of a rant, but I'd far prefer us all focusing on open standards like HIP, SYCL, OpenCL, etc.
It's unbearable that most LLM inference happens on closed H/W, closed drivers, and closed SDKs.
mistercow 11 hours ago
On the SDK front, you tried just having an agent reimplement the model you're interested in and just use the weights? I've taken to treating off the shelf implementations as reference implementations anyway, because I can often squeeze out significantly better performance for my configuration and use case by having Codex hammer at it for a few hours.
sroussey 10 hours ago
Hugging face is working on something like this where well known models get fused into a single implementation.
drivebyhooting 10 hours ago
Could you share the prompts and workflow? I’ve tried this too, but with mixed success when it comes to creating custom tile kernels.
I would really appreciate your input!
mistercow 10 hours ago
mschuetz 9 hours ago
The problem with the open standards is that their dev UX is absolutely horrible. You can't neglect usability, and then be surprised that there are no users.
the__alchemist 8 hours ago
This is at the core of the matter for me, and my knowledge is too weak to understand why this is the case. I don't enjoy the idea of relying on Nvidia's stack for GPU compute, but the alternatives I've tried (e.g. Vulkan compute) are higher friction to use. I am trying to reconcile why; Nvidia shouldn't have this moat.
My software is labeled "CPU only unless using an nVidia GPU". I would prefer to strikethrough "nVidia". Incidentally, this means no more Mac support.
swerner 7 hours ago
christopher8827 an hour ago
yeh - exactly.
Sucks like important libraries like Alphafold are locked into CUDA. Its ridiculous for researchers.
HeavyStorm 9 hours ago
Commercially it's better to have AMD support CUDA which can help break Nvidia soft monopoly.
bigyabai 11 hours ago
I would too, but sadly that's Khronos' job to organize, and they've had trouble getting American vendors to work together.
It's likely that CUDA will continue dominating until they put aside their differences. The current MLX/MPS/ROCm ecosystems are too fractured to threaten Nvidia.
boredatoms 11 hours ago
Maybe a not-Khronos org should try
high_na_euv 11 hours ago
Wdym American Vendors?
Intel uses SPIRV iirc
my123 10 hours ago
bigyabai 11 hours ago
kiicia 9 hours ago
it will only get worse now when hughingface was bough by nvidia, not immediately, but in span of year or years...
anon291 8 hours ago
It's impossible to have an open standard. Hardware accelerators are nothing alike and have different perf characteristics. Each kernel is tuned to the hardware. The idea of writing a performant kernel in opencl is a fantasy.
Source: worked at a bunch of accelerator companies in the kernels or equivalent team. They're nothing alike.
swerner 8 hours ago
For a performant portable language, we’d have to go to a higher level where you describe what to do and leave the how to do to the compiler. It would then need to be able to adjust memory layout, access patterns, data type choice to the underlying hardware. I’m not sure if this is possible to do reliably - the closest we have right now may in fact be highly detailed plain English descriptions of the algorithms fed to an LLM prompted to produce assembly.
anon291 7 hours ago
swerner 11 hours ago
AI will take down Nvidia’s moat. When it becomes trivial to translate CUDA/PTX to HIP, SYCL or Metal, CUDA is no longer the moat, it becomes the intermediate representation.
larodi 9 hours ago
trivial to translate (or transpile) - okay. trivial to understand the result - not so much. trivial to then evolve it - hm... perhaps a different story. still, it seems very likely now, that such "quick rewrites" are viable, not sure if an open approach to them is viable. a newly born open project that was LLM-derived, and not by a credible author, which spans hundreds of files no human eye has ever looked at, can only work for a closed organization, but will never be trusted by the general audience... just like that.
threatripper 2 hours ago
I don't see a hard reason. If it works it works. No hard need for a good, universal, and long lasting solution. At some point you just stack slop on top of slop and it works for your use case - and if it doesn't you'll slop it out yourself.
saagarjha 8 minutes ago
bayindirh 10 hours ago
> When it becomes trivial to translate CUDA/PTX to HIP,...
ZLUDA is already doing that, no?
swerner 9 hours ago
I don't think we're at a point yet where anyone would trust ZLUDA enough to ship commercial products that rely on it. I would be delighted though, if anyone can prove me wrong.
bayindirh 9 hours ago
Keyframe 10 hours ago
yeah yeah, "when" an often keyword with AI it seems. As Mr. E. Nigma put it - what always comes but never arrives? Meanwhile the moat deepens and it's build on inertia and laziness and Nvidia knows this really REALLY well.
swerner 9 hours ago
Oh, absolutely. Nvidia is the modern day “nobody gets fired for buying IBM”. The reason we’re still using Unix is not because it’s the best, but because it had to much inertia to let any alternative become its successor. Similarly, C and HTML are maybe the most terrible yet extremely useful languages we have.
mathisfun123 10 hours ago
i swear people who are outsiders here have only clickbait takes; if you've never had to ship GPU code professionally you should just not comment on these things.
the source language has never been the moat. Nvidia sells to hyperscalers. Hyperscalers have armies of kernel authors who have no issue translating shaders by hand (or now with claude). Nvidia's moat is (and will remain for the foreseeable future) the entire stack. you cannot fathom the pain and misery of working on literally any other stack. if you've never debugged a GPU synchronization error or kernel panic due to some GPU firmware bug or fought absolute shit profilers hunting for perf you really have no idea what you're talking about.
EDIT: i can't believe this really requires saying but graphics and compute are not the same domain at all. if you work in graphics for GPU but not compute then you are still way out of your depth commenting. to wit: graphics people do not (and cannot) write CUDA kernels/shaders.
swerner 10 hours ago
If that is your standard, I do have an idea what I’m talking about.
mathisfun123 10 hours ago
swerner 9 hours ago
Most modern graphics is compute. Pixar, Dreamworks, Sony, etc do not use Vulkan to render their movies. It’s CPUs or CUDA.
“graphics people do not (and cannot) write CUDA kernels/shaders” is just not true at all. All it would take to verify that would be things like reading the introduction of the OptiX documentation, a small sample of SIGGRAPH GPU papers or the Blender/Cycles source code.
vivzkestrel 12 minutes ago
stupid question
- what exactly is CUDA?
- is it patented hardware tech that AMD cannot replicate?
- is it a library that ll also work on AMD while it currently works on nvidia
- can AMD cook an equivalent of CUDA, if so why havent they done it?
triwats 5 hours ago
Interesting option for CDNA architecture chips. I wonder if this moves to an open standard?
AMD GPUs build for AI specs for reference: https://flopper.io/gpus?vendor=AMD&page=1
lulzx 11 hours ago
I made cuda-metal btw (for mac kek), https://github.com/lulzx/cuda-metal
sroussey 10 hours ago
What models can it run?
KennyBlanken 6 hours ago
To save everyone a click: No cuDNN and based off an ancient version of ROCm for windows (7.1 has been out for ages, 7.2 is current.)
Nurysso 9 hours ago
man i can't say how much i used to like cuda when i had a nvidia gpu it made ml so much fun and on amd its a war especially on rdna 2 cards which i have. hopefully one day we will be able to properly translate cuda for its amd counter parts
system2 12 hours ago
I wish there were a way to use RDNA1 cards with CUDA for AMD. My 5700XTs are sitting in a drawer.
monster_truck 11 hours ago
RDNA1 isn't good for a whole lot, even flagship RDNA2 cards are a stretch for many things. The lack of WMMA/matrix multiply/BF16 is too severe of a penalty.
The FP16 throughput on RDNA1 is both shader reliant and requires everything to be packed first. Even with 2 or 4 or 1000 cards, you would be consuming all of the available memory and memory bandwidth just packing and unpacking values, and if you really want to dump a hundred billion tokens into making it work anyways, you're only going to find out that even if you bother to sit there ferrying packed values to ram or disk before then issuing the instructions, paying that already severe penalty again when the values then have to be unpacked is so steep of a cost that the 256 BF16 flops/cu/clock's effective throughput is outright lower than simply doing it on a Zen 2 processor. You also don't have INT8 (or really INT4) on RDNA1 so the other RNS/CRT tricks aren't viable.
Sadly RDNA1's VCN2 also lacks actually good x264 bframe encoding support, or even P010 for 10 bit color, so what I'm saying is you should sell them. Used Radeon VII's are like $260, you'll go a lot further with those especially if you throw in a 7900XTX, and then augment that further with a 9070 CRE (you only want it for its int8 cores), and of course 128GB of ram.
E: And sure, that's 3, or ideally 4 GPUs, and a good bit of extra work. But that gets you up to more than halfway to the naive performance of a $15,000 MI300x in a surprising amount of cases, with additional strengths that it lacks. For far less than half of the cost
system2 11 hours ago
I agree, I should've sold them last year when they hit $500 each. I have like 8 of them from my mining days. What a silly mistake I made.
dracotomes 9 hours ago
latchkey 9 hours ago
chiassedu80 14 hours ago
CUDA for AMD on Windows
I’ve been working on a Windows setup that lets CUDA-targeted applications run on AMD GPUs using ZLUDA + ROCm/HIP.
Repo: https://github.com/Speedstu/CUDA-for-AMD-Windows
So far, it has only been tested on my RX 9060 XT (gfx1200), where I’ve used it with CUDA-enabled LibTorch workloads, including long ai training and use.
I also added a GPU scanner / auto-detection system that detects:
AMD GPU model
gfxXXXX architecture
ROCm/HIP installation
driver info
whether the GPU has already been validated by the project
Example:
RX 9060 XT → gfx1200 → RDNA4 → HIP detected → validated
The goal now is to test it on more hardware, especially RX 6000 / 7000 / 9000 cards.
If you have an AMD GPU on Windows and want to try it, I’d really appreciate compatibility reports working or broken. There’s a dedicated GPU compatibility issue template in the repo.
If this is useful to you, a star would also help the project get more testers.
nine_k 11 hours ago
(As a side note, I love the name ZLUDA; it very aptly means "delusion" or "deception" in Polish.)
swerner 8 hours ago
TIL, I didn’t know that. I always assumed it came from level zero”, the Intel computer layer that ZLUDA was translating to before its developer was hired by AMD to target HIP.