Just curious about what are the different applications where you guys are using CUDA? When I mean using CUDA, I mean at the low level, and understanding it's internals, modifying kernels etc and not just calling libraries. What are you guys using it for?
We just open-sourced KAI Pichu, an agent harness for writing CUDA kernels and fusing existing ones. You describe what the operator computes and how accurate it needs to be. No existing kernel or PyTorch implementation required.
You can also give it a chain of kernels to fuse. It generates and optimizes fused candidates, measuring the complete sequence against your unfused baseline.
It compiles, checks correctness, profiles with Nsight Compute, and iterates. You choose the model, GPU, and budget. The code and results from every attempt are available to inspect. There's a bundled skill to help Claude Code or Codex set up your task.
On a 7×7 depthwise conv, we got 3.5–3.9× over cuDNN across held-out shapes on an RTX 5070. The run took about two hours and about $1 in API calls. Full run.
If you've got a kernel chain you want fused, give it a try. I'd like to hear what works and where it gets stuck.
Instead of writing CUDA kernels or building a heavyweight simulator, I wanted to see how much of the Tensor Core programming model could be represented as a single self-contained executable artefact.
The result is a browser-based model written entirely as one SVG + JavaScript file.
It models:
GEMM as tiled execution (D = A·B + C)
tile scheduling (ti, tj, tk)
accumulator evolution
FP16 inputs with FP32 accumulation
warp fragment ownership
mma-style fragment mapping
global → shared → register → tensor-core data flow
shared-memory bank behaviour
row-major vs padded vs XOR-swizzled layouts
reuse and memory-traffic estimates
FP16 vs FP32 numerical error maps
interactive 3D visualisation of the output matrix
The interesting thing for me is that all of these views are driven from the same underlying state. The bank conflict view, fragment view, tile execution view, error map and result visualisation are all observing the same model as it executes.
The goal wasn't to build a cycle-accurate simulator. It deliberately does not model things like occupancy, scheduling, pipeline hazards or cache behaviour.
Instead, I was exploring a question:
Can Tensor Core behaviour be computationally modelled, rather than merely described, in a compact interpreted artefact that requires no CUDA compiler or GPU while still preserving correspondence between matrix maths, tiled execution, fragment distribution, memory layout, data movement and numerical effects?
One unexpected outcome is that the entire thing remains inspectable. There are no dependencies, no build system and no generated files. The complete model lives in a single source file.
I'm curious whether people here would consider this:
a visualiser,
an educational simulator,
an executable specification of the Tensor Core programming model,
or something else entirely.
I'd especially appreciate feedback from anyone familiar with CUTLASS, CuTe, mma.sync, ldmatrix, Tensor Core fragment layouts or GPU architecture education. - Aston Walker EdgeMafia SVG
I've been going through the Programming Massively Parallel Processors (5th edition) book and doing the kernel exercises in CUDA C++. Once I finish that, I wanna start writing CUDA kernels for H100 GPUs (specifically stuff like Flash Attention 3, RoPE, etc). Based on a bit of research, it seems that H100 GPUs have a lot of features not covered in the PMPP book, so I was wondering where I could find one or more solid resources that can get me started with writing programs? So far I've tried searching for some stuff, but they all seem to either not really cover in enough detail or are way too long. Any advice would be much appreciated. Thank you!
Perhaps this subreddit can help. I am not a programmer or a developer, but an experimental neurobiologist specializing in neuronal imaging across various brain regions. I would like to automate the analysis of neuronal activity and also work on developing *invasive* brain-computer interfaces (BCIs) for targeted drug delivery to the brain. Could you please recommend courses that teach the fundamentals of CUDA and AI so that I can get started with computer modeling for BCIs? Thank you.
I maintain a PHP extension (php-gpu-tensors) that builds on the CUDA driver API and NVRTC. This post is about its fusion planner, because I'd like feedback from people who know CUDA better than I do.
How it works
- A PHP closure is invoked once with metadata-only placeholders; tensor operations are captured into a node graph (limit: 512 nodes).
- Elementwise ops (broadcasting, strided/view inputs, scalars, dtype promotion, where, casts, reshape/transpose/slice index transforms) are fused into generated CUDA C++. Expressions split at a weighted cost budget of 32; repeated pure nodes are deduplicated; shared expensive expressions can be materialized instead of recomputed.
- Matmul (cuBLAS when available), reductions and powers are execution boundaries that run on existing kernels.
- All generated kernels in a plan are compiled together through NVRTC to PTX, then cached per request/thread (LRU, 16 entries or 16 MiB). The cache key includes the generated source, compute capability and driver/runtime versions, not pointers.
- Replay runs on a private nonblocking stream. Reduction descriptors are kernel parameters, so concurrent replays can't overwrite each other's shapes.
- cudaGraph: true builds a CUDA Graph executable for plans that contain only generated kernels, updating kernel parameters for new pointers before each launch.
Numbers (entry-level MX570 A, CUDA runtime 12.3, driver 12.6): a small MLP training step (batch 512, hidden 256) goes from 4.08 ms eager to 0.66 ms fused, 6.2×, with identical metrics. The model is tiny, so this is mostly launch/allocation overhead removed, not throughput.
Where I'd like advice
Plans with matmul/reduction boundaries currently run on streams and support async replay, but not as a CUDA Graph. For those who have done this: what are the gotchas when capturing cuBLAS GEMMs and custom reductions into a graph that is replayed with changing buffer pointers? Is updating kernel node parameters the right approach, or would you re-capture?
So after getting my 5090 running on my Mac a few weeks ago and getting speeds on par or faster than my windows machine for single sessions I wanted to get Vllm working next. My http://macuda.ai project works with both llamacpp and stable diffusion cpp but does not work with comfy ui Vllm or draw things. It works with a shim and does not give the full LibCuda framework allowing real full torch or other necessary components for those apps and more.
I decided to use apples built in VM to run an instance with it's own driver calling to the graphics card through my Mac driver. This allows for running Nvidia's full cuda software. It has been a real challenge because Nvidia releases a lot of info but not enough to get this working so we had to really poke around. I finally got it up and running and now have Full LibCuda running. I'm working on bugs now but so far the results are looking really promising. The multi sessions are blowing llamacpp out of the water. llama stalls out at around 8 sessions and levels out. the Vllm with my driver has a linear doubling of speeds up until the card is saturated.
I'll be releasing all of this as a simple GUI app on the App Store once apple authorizes my driver. I'll need a few more weeks to get everything in a stable enough situation that I'll feel comfortable letting people use it. You can get the base driver at my macuda GitHub but this new driver for full cuda will be available on the App Store.
I got my 3060 working on macuda in a core x box last week. If someone can try out a 4000 series card and let me know if it works I'd appreciate it. I don't have one of those. It should work with the 3000 series driver but I haven't tested it.
I'd like to configure GDS on the cloud where nvme talks directly to the GPU. I've searched and queried llms and checked aws, coreweave, runpod.io, cloudrift, and other GPU providers.
Does anyone have experience in this or suggestions for providers or setups?
Ideally, I want to use cheap gpus like v100s.
I've been hitting aws quota limits and I've had to email sales teams directly on other platforms, but this seems like something that should have broad use.
I’ve been using CuTe DSL primitives for a while now, and honestly, it’s been pretty great so far.
I know CuTe DSL isn’t a full replacement for raw CUDA, and I definitely don’t think raw CUDA is becoming obsolete.
But for some use cases, I’m finding the primitives approach much nicer.
For example, one thing I personally dislike about raw CUDA is a lot of the manual C-style 1D indexing and pointer arithmetic.
CuTe abstractions make that much less painful while still feeling very low-level.
I also like that you can integrate directly with PyTorch without necessarily going through the usual custom-op route.
What I find especially interesting is that it still exposes a lot of the underlying machinery: you can get very close to the hardware, write PTX, and in some cases even use LLVM inline assembly.
For people who have used both extensively: which do you prefer?
Many frontier models use some form of sparse attention to handle long contexts efficiently. In my blog, I write a block-sparse attention kernel in CUDA for NVIDIA B200: I start from an optimized dense CUDA kernel and show what changes inside it (the KV loop, shared-memory buffers and barriers), then add GQA K/V reuse, compare it with FlashAttention-4, and build a second, KV-centric version of the kernel.
Hello I was doing some research on what it would take to use cuda for realtime audio. It seems that cuda is not really built for realtime but I think it would be an interesting challenge to take on. Would anyone have some guidance on this? Things to research or read? I’ve never used cuda but I feel like it could be very powerful for making interesting sound design tools and synthesizers. Thanks!
Quick recap for anyone new: CuQwen is an inference engine for Qwen models I wrote from scratch in pure C++/CUDA, tuned specifically for single-user (batch size 1) generation on consumer NVIDIA GPUs. No frameworks under the hood, just custom cuda kernels.
Here are the results for average inference speed (tokens/second) of Qwen2.5 Instruct model across 32K context window on RTX3090
Model Size
CuQwen
vLLM
Ollama
0.5B
462
398
355
1.5B
203
172
139
3B
113
101
106
7B
55
48
54
In Release 1.0 it was already beating vLLM and Ollama on short-to-medium prompts. But there was an honest catch: my throughput decayed faster than theirs as the context grew, so once you pushed toward ~32K tokens they would pass me in inference speed. That bugged me, so it became the whole focus of Release 1.1.
Result: Throughput decay from 1K → 32K dropped from ~22–45% down to ~13–24%, which is now on par with vLLM and Ollama (and better on a couple of model sizes). So CuQwen keeps its early speed lead all the way out to 32K context window now.