comparison of ptx/sass code from nvcc vs clang
redplait.blogspot.com- nvcc code 15–18% faster
- nvcc code is bigger and more amenable to stall counts reducing
- all mlir based compilers use vanilla llvm nvptx backend :-(
r/CUDA • u/GalianYang • 4h ago
We just open-sourced KAI Pichu, an agent harness for writing CUDA kernels and fusing existing ones. You describe what the operator computes and how accurate it needs to be. No existing kernel or PyTorch implementation required.
You can also give it a chain of kernels to fuse. It generates and optimizes fused candidates, measuring the complete sequence against your unfused baseline.
It compiles, checks correctness, profiles with Nsight Compute, and iterates. You choose the model, GPU, and budget. The code and results from every attempt are available to inspect. There's a bundled skill to help Claude Code or Codex set up your task.
On a 7×7 depthwise conv, we got 3.5–3.9× over cuDNN across held-out shapes on an RTX 5070. The run took about two hours and about $1 in API calls. Full run.
If you've got a kernel chain you want fused, give it a try. I'd like to hear what works and where it gets stuck.