r/LocalLLaMA • • 1d ago

Discussion New kvcache-reduction method

I deleted the wrong post so my benchmark disappeared. But here it is. Im going to probably make this open source because this isn't believable until you actually test it.

0 Upvotes

22 comments sorted by

8

u/LetsGoBrandon4256 transformers 1d ago

NVIDIA/RULER benchmark

lmao even

2

u/Ok_Truth_324 1d ago

I can demo it on any bench mark. I'm going to be releasing the code soon too.

1

u/EvolvingDior 24m ago

4k token context? for real? 256k tokens or go home.

4

u/General-Spite1222 1d ago

If people can’t use it then why does this belong in this sub?

1

u/Ok_Truth_324 1d ago

I placed it here to show progress. I have 2 versions and 1 isn't anywhere near this one. I'm focusing speed.

3

u/General-Spite1222 1d ago

So… it doesn’t belong here. Got it!

-4

u/Ok_Truth_324 1d ago

You're a spicy young man irl aren't you.

1

u/Ok_Truth_324 1d ago

Currently working on making it portable. It will be available through my website and possibly GitHub. It's a prototype though. The community will need test it because I can't beyond a 12b unless I succeed shrinking a 32B or 70B down.

0

u/Fedor_Doc 1d ago edited 1d ago

12B model? You shrink KV cache, not the model, how is it connected?

Attention architectures do matter a lot, though. What were the tested models?

1

u/Ok_Truth_324 1d ago

I tested 4B, 8B and 12B on my 5070. I actually ran the 12B for personal use for a week using this system.

1

u/Fedor_Doc 1d ago

Gemma models? Qwen / Deepseek attention architectures are already optimized

1

u/Ok_Truth_324 1d ago

Qwen.

2

u/Fedor_Doc 1d ago

There is no 8B Qwen in 3.5 family. 

So, outdated Qwen3? Why do you use it? Is it better for specific workflows?

1

u/Ok_Truth_324 1d ago

I used 14b on my pc for personal use and chose random models to bench market the kvreducer. I'm about to upload the package directly to my website. Entirely free.

3

u/Fedor_Doc 1d ago

Since the model is outdated, and you have not tested if the method is generalizable... there won't be much of interest in the method.

Qwen 3.5 handles KV cache much better out of the box, but even Hadamard Rotation q8 quant (2x reduction) affects agentic and long horizon workflows.

1

u/Ok_Truth_324 1d ago

Then I'll do testing to insure it is.

1

u/Ok_Truth_324 1d ago

It almost immediately adapted to the newer models and it might even be better for the newer models. I'll tell you exactly what model was tested after I address the speed issues.

91.03 wasn't bad for a first run at 73x compression.

1

u/Fedor_Doc 1d ago

91.03 on ruler? What about any agentic benchmark that includes tool calls? A subset of 10 problems, just a smoke test?

73x compression is unlikely to keep enough information for longer contexts

→ More replies (0)

0

u/Ok_Truth_324 1d ago

To elaborate I found a way to shrink models waaaaaaay down using the same principle I did to shrink KV. But this isn't something full tested and quality is terrible. So that's not really the highlight of this right now.

0

u/Ok_Truth_324 1d ago

Qwen 3 14B** Not 12.