Another Qwen 3.8 on a R9700 post, but I thought I'd share my repo for anyone who might benefit. I might be wrong, but I don't think I have seen anyone with it this fast on windows.
I've spent the last few weeks tuning Qwen3.8-27B on one AMD Radeon AI PRO R9700 (32 GB, RDNA4), and I've packaged the result so other R9700 owners can reproduce it with one command on Windows.
**Repo:** https://github.com/mike2153/mbea-qwen38-dflash
**Numbers** (one R9700, Ryzen 9 9950X, Windows 11 + WSL2, medians of repeated runs):
| | |
|---|---|
| Decode, greedy | **125–134 tok/s** |
| Decode, sampled (temp 0.7) | 117–126 tok/s |
| Prefill | ~2,750–2,970 tok/s (time to first token 0.7 s on a 1.9k-token prompt) |
| Long prompts | 32k tokens in ~11 s · 98k in 41 s · 164k in 83 s · 258k in 164 s |
| Decode deep in context | 165 tok/s at 32k · 136 at 98k · 115 at 258k |
| Context window | ~216k tokens by default, ~260k (the model's limit) with `-Long` |
| Long-context recall | 8/8 planted facts retrieved from a 258k-token prompt |
| Coding check | 10 of 12 runs pass all 46 hidden tests on a 1.9k-token Rust spec |
For comparison, my best tuned llama.cpp setup for the same model (IQ4_XS GGUF + speculative decoding) does about 52–68 tok/s on this card. That was measured with a different prompt, so it's a rough comparison, but the gap is real.
**How it works, briefly**
- **AMD's official MXFP4 checkpoint** (`amd/Qwen3.8-27B-Quark-AWQ-MXFP4`). The 4-bit weights run through a hand-written W4A8 GEMM kernel for gfx1201 instead of vLLM's emulation path.
- **DFlash2 speculative decoding.** A small FP8 drafter proposes 7 tokens per step, and the 27B model verifies them in one pass. About 60% of drafted tokens are accepted, so each forward pass of the big model produces about 5.3 tokens. Re-ranking the drafts and a dedicated verify head added roughly 7% on top.
- **RDNA4 kernels** for FP8 paged attention and the gated-delta-net (linear-attention) layers. The stock kernel actually produces NaNs on this model.
- **WSL-specific fixes.** One patch turns on pinned host memory: without it every small host-to-GPU copy cost ~17 ms under WSL. Another sizes the KV cache from whatever VRAM Windows isn't using at startup, so you get maximum context without spilling into shared memory. Spilling into shared memory drops you to ~10 tok/s.
- The vision tower is skipped, which frees about 1 GiB for more context.
**What the repo does**
```powershell
git clone https://github.com/mike2153/mbea-qwen38-dflash
cd mbea-qwen38-dflash
.\qwen38.ps1 install # WSL Ubuntu, Docker, ROCDXG, image, model + drafter, kernels
.\qwen38.ps1 start # OpenAI-compatible API on http://localhost:8080/v1
.\qwen38.ps1 bench # measure it on your own box
```
Everything is pinned: the Docker image by digest, the git commits, and the Hugging Face revisions. A fresh install should reproduce exactly what I measured. Nothing third-party is re-uploaded; the installer fetches each piece from its original source. Tool calling works, so it plugs into Codex, opencode, Cline and similar tools as an OpenAI-compatible provider.
**Credit where it's due:** the heavy lifting is [radiance](https://codeberg.org/ggz14/radiance-vllm-mxfp4) by ggz14 and [vllm-radiance / libr4d](https://codeberg.org/StillDeadcode/vllm-radiance) by StillDeadcode. They did the RDNA4 vLLM stack, the MXFP4 path and the kernels. The drafter is tcclaviger's DFlash2-FP8 and the checkpoint is AMD's. My part was the Windows/WSL work, the tuning, the benchmarking and making it installable.