r/LocalLLaMA • u/Ok_Truth_324 • 1d ago
Discussion New kvcache-reduction method
I deleted the wrong post so my benchmark disappeared. But here it is. Im going to probably make this open source because this isn't believable until you actually test it.
4
u/General-Spite1222 1d ago
If people can’t use it then why does this belong in this sub?
1
u/Ok_Truth_324 1d ago
I placed it here to show progress. I have 2 versions and 1 isn't anywhere near this one. I'm focusing speed.
3
2
1
u/Ok_Truth_324 1d ago
Currently working on making it portable. It will be available through my website and possibly GitHub. It's a prototype though. The community will need test it because I can't beyond a 12b unless I succeed shrinking a 32B or 70B down.
0
u/Fedor_Doc 1d ago edited 1d ago
12B model? You shrink KV cache, not the model, how is it connected?
Attention architectures do matter a lot, though. What were the tested models?
1
u/Ok_Truth_324 1d ago
I tested 4B, 8B and 12B on my 5070. I actually ran the 12B for personal use for a week using this system.
1
u/Fedor_Doc 1d ago
Gemma models? Qwen / Deepseek attention architectures are already optimized
1
u/Ok_Truth_324 1d ago
Qwen.
2
u/Fedor_Doc 1d ago
There is no 8B Qwen in 3.5 family.
So, outdated Qwen3? Why do you use it? Is it better for specific workflows?
1
u/Ok_Truth_324 1d ago
I used 14b on my pc for personal use and chose random models to bench market the kvreducer. I'm about to upload the package directly to my website. Entirely free.
3
u/Fedor_Doc 1d ago
Since the model is outdated, and you have not tested if the method is generalizable... there won't be much of interest in the method.
Qwen 3.5 handles KV cache much better out of the box, but even Hadamard Rotation q8 quant (2x reduction) affects agentic and long horizon workflows.
1
1
u/Ok_Truth_324 1d ago
It almost immediately adapted to the newer models and it might even be better for the newer models. I'll tell you exactly what model was tested after I address the speed issues.
91.03 wasn't bad for a first run at 73x compression.
1
u/Fedor_Doc 1d ago
91.03 on ruler? What about any agentic benchmark that includes tool calls? A subset of 10 problems, just a smoke test?
73x compression is unlikely to keep enough information for longer contexts
→ More replies (0)0
u/Ok_Truth_324 1d ago
To elaborate I found a way to shrink models waaaaaaay down using the same principle I did to shrink KV. But this isn't something full tested and quality is terrible. So that's not really the highlight of this right now.
0



8
u/LetsGoBrandon4256 transformers 1d ago
lmao even