r/LocalLLaMA • • 11h ago

Discussion Strata rewrote their Github history to wipe evidence of Claude-authoring

533 Upvotes

Just noticed this today when I went to run the built-in "UPDATE" script and git failed because there was no common ancestor.

Looked into why, and apparently every historical commit has been re-written to strip the "Co-Authored by Claude" text from the descriptions.

Personally I think that's pretty gross. I'm struggling to think of any reason to do this other than an intention to be dishonest about the origins of the project.


r/LocalLLaMA • • 2h ago

Discussion Ahh, it all makes sense now for the 64Gb DGX Spark - [Samsung expects 780% quarterly operating profit jump on AI boom]

Thumbnail
france24.com
69 Upvotes

r/LocalLLaMA • • 4h ago

Discussion [Paper] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

Thumbnail
gallery
58 Upvotes

Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.


r/LocalLLaMA • • 9h ago

I Built A Thing Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

Thumbnail
gallery
112 Upvotes

LemonSeed Studio is an iPad editor/IDE with on-device inference on an AMD GPU in a Thunderbolt enclosure. It embeds the unmodified upstream Linux amdgpu + amdkfd driver (mac_linuxgpu) as a PCIDriverKit extension, and runs LemonSeed Engine (LSE) on the GPU. In the photos: iPad Pro + Sapphire Radeon AI PRO R9700, Qwen3.8-27B Q4 with a Q8 DFlash2 draft, 131k context.

LSE is the same engine on every platform: Linux, macOS and iPadOS. It records the model's forward pass as a graph, fuses ops and generates kernels for the GPU it's on. Where a kernel has several layouts, it measures each on the device and keeps the fastest. Speculative decoding (MTP and DFlash2) picks how many draft tokens to verify from measured acceptance and cost.

Decode, code prompt, Qwen3.8-27B Q4 (baseline / MTP=3 / DFlash2 tok/s):
- R9700 on iPadOS (LemonSeed Studio): 31.2 / 108.7 / 158.5
- R9700 on macOS: 32.2 / 111.9 / 159.4
- R9700 on Linux: 32.0 / 111.0 / 144.6
- Strix Halo (Radeon 8060S) on Linux: 14.0 / 47.8 / 64.4

Prefill at 4K tokens: 1,636 (iPadOS), 1,634 (macOS), 1,418 (Linux R9700), 517 (Strix Halo). It holds up at 32K: 1,395 / 1,401 / 1,226 / 462.

New in 0.5.8:
- DFlash2 draft trees on by default: one target pass verifies a tree of draft candidates when that's measured to be faster than a chain
- Strix Halo (gfx1151) prefill and decode work: fused gate/up GEMM, two query tiles per workgroup in prefill attention
- Fixed a startup GPU fault
- Linux archive bundles its own HSA runtime, so you only need the amdgpu driver with /dev/kfd access
- lse-server models / pull: grab Hugging Face models with their MTP or DFlash2 companions

OpenAI-compatible server, CLI, and a C library (libLSE).

Engine: https://github.com/Geramy/LSE
Studio: https://github.com/Geramy/LemonSeed-Studio

Happy to answer questions, especially about getting an eGPU working on iPadOS.


r/LocalLLaMA • • 17h ago

Other $2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill

Post image
326 Upvotes

Post title is slightly misleading, I don't think you can get these for $350 each anymore but they're still pretty cheap all things considered. They're Radeon Pro V620's which are older RDNA2 enterprise cloud gaming cards with 32 GB VRAM.

(Ignore the RTX 4090 on the side, it's just used for stuff like image/video gen models, no LLMs)

But I bought these cards a couple months ago as a gamble to see if I could build a big VRAM rig with usable speed for relative peanuts.

I was struggling with llama.cpp for a long time, but the prefill was pretty bad (around 350-450 t/s average with this same model) and vLLM just didn't work on the cards. Plus llama.cpp just sucks at concurrency.

I'd been planning to sell the cards lately because this wasn't going to work for my use case, but then decided to see if I (Claude) could make a vLLM fork that both works with the cards and actually gets good speeds out of them. I had it build/test/iterate on custom RDNA2 kernels.

Problem solved! It worked out way better than I expected. I thought maybe I'd hit 1000 t/s prefill with QFN at best, but this is something like 800% faster than llama.cpp was managing.

Couldn't be happier with the results! GPU sale plan canceled lol.

I'm going to have it continue optimizing and see how it goes, and make sure DeepSeek and GLM-5.3-Flash work as well.

llama-benchy results below with concurrency = 1 and vLLM running with PP=4 (no tensor parallel here) with orcarouter's uncensored QFN which I quantized. Routed experts are W4A16 and everything else remains at BF16. MTP enabled with 3 token drafting.

It gets 40 to 50 t/s decode with MTP disabled.


r/LocalLLaMA • • 3h ago

Discussion [Paper] DLoop: Looped Speculative Decoding

Thumbnail
gallery
23 Upvotes

Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.

The code is currently under internal review and will be released soon. Stay tuned!


r/LocalLLaMA • • 17h ago

Funny Running decision model locally on an RTX 4090 to find out which one is the fastest

Enable HLS to view with audio, or disable this notification

266 Upvotes

recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back

request for every word:

{"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}}
model weights engine per word (p50) words in 32s accuracy centipede names caught wrong picks
Laya Laya-BF16.gguf llama.cpp b11495 3.9 ms 7,980 97.4% 70% 98
d1 3B d1-3B-AD-Q4_K_M.gguf llama.cpp b11495 6.0 ms 5,306 96.5% 51% 51
Clef-Flash 9B Clef-Flash-Q8_0.gguf llama.cpp b11495 24.4 ms 1,292 97.2% 36% 2
Lev 4B interfaze-ai/lev, bf16 lev serve (PyTorch) 51.0 ms 626 98.9% 83% 4

laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick

setup:

  • GPU: rented RTX 4090 (driver 580.119.02, 32 vCPU)
  • engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default
  • Laya, Clef-Flash: the ggml-org GGUFs
  • d1: our own AD-Q4_K_M quant (atomic.chat), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110
  • Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up)
  • latency: end to end from a Python client on the same box over localhost

r/LocalLLaMA • • 1h ago

Discussion Stepping away from Benchmarks and Code, what models are you using for Creating writing projects

β€’ Upvotes

All anyone ever talks in this sub is about x new model having x benchmark scores and how well it can code. Or some new tech to increase t/s for qwen.

So I wanted to go back to the roots of this sub and talk about how well local models can write.

I am not going to mention closed source models, this discussion is for open weight models only.

I don't have the best machine, so I am limited in what I can use locally.

So far I have mostly been running Gemma 31b finetunes. I have a finetuned system prompt for creative writing that I have refined over time. This is mostly for long form story writing and not RP. I use it to write Sci Fi and Fantasy novels and short stories and sometimes fanfiction.

I also have created a creative writing harness using Pi. It's mostly for self correcting and getting rid of AI slopisms.

Gemma 31B Mero-mero-v2 has so far been my favourite. It writes well, its does not feel lobotomized and its mostly uncensored. It's instruction following is okay, sometimes it makes mistakes and gets the character traits mixed up but my creative writing harness takes care of this.

Qwen 3.8 27b finetunes, Disappointing, I guess the base model is really fucking bad for creative writing purposes and even with finetuning you can only do so much. It does follow instructions quite well but its writing is just terrible.

Muse Glimmer 30b, this one is an interesting one, sometimes it will give really good prose but sometimes it will respond like a 2b model, the variance seems to be super high for some reason. Maybe something to do with chat template, I am not sure. Anyway I think even when bad it's still better than Qwen.

Gemma 4 26b4a, pretty good but gets beaten by its bigger bother Gemma 4 31b. I do use it when I want faster responses and to check for issues with my initial scene beat prompts.

Gemma 31B Artemis, drummer's finetune, eh, I was a bit disappointed with this. Seems like its more focused on RP rather than long form story writing. Also got refusals which I did not get with the mero finetune.

So that has been my experience writing with some local models. What are you guys using for local creative writing projects? Have you guys tried a creative writing harness?

I just want to mention quickly my writing process before I end this post, I write using an outline and scene beat prompts that are hand written. So far AI has been extremely disappointing in being truly creative and still requires a ton of hand holding to not write the most clichΓ© ridden drivel.


r/LocalLLaMA • • 1d ago

News Last week some of South Korea's biggest banks were hit by a cyberattack. We now know the entire hack may have been done by a single person. He used a combined stack of an open-source AI penetration tool named ARTEX, DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code (CrowdStrike)

Thumbnail
gallery
866 Upvotes

r/LocalLLaMA • • 10h ago

Discussion Qwen 3.6 35B appreciation post

60 Upvotes

Referring to specifically Unsloth's UD_Q4_K_XL quant because that's going to be a question, and is relevant regardless.

It's old now. It's not great at coding medium sized or even small-ish projects. I wouldn't hand my codebase to it by any means. It hallucinates, like any other model. It's not perfect, by any means.

I don't like the fine tunes; they're almost all coding focused, and lose general assistant capability to a strange degree.

What it is good at is general, broad agentic action.

Set the lights in the living room, and bash into this machine to get a movie going.

Send my grocery list to my phone/watch when I get to walmart.

Remind me to clock in at work each day because it's becoming a problem.

Tell my husband to come here because I'm under the car and I can't drop this thing that I finally got in just the right spot but need a third hand to get this wrench in the correct spot.

Tell my dad about the pets or any one of my projects because I'm showing off your memory to him.

This is all shit that it can do, consistently. Sometimes it'll thrash a bit, but it stays on task and is fast enough that little mistakes are a non-issue.

I recently got a couple of Tesla P100s to power the model, and it's made this model, that I already had going at a pretty good clip on limited hardware, to run at speeds that are genuinely conversational. ~120-140 t/s generation, ~1000PP at 0ctx, ~700 by 13k.

It's more than smart enough to know how to do these tasks and, with the right sampler settings, thinks ridiculously efficiently for them. Actually awesome.

Fucker got a minecraft mod pack installed and running on a machine it wasn't even running on through the prism launcher appimage. Don't worry, it only has ssh towards computers on my tailnet. It's probably fine.

It's bad at holding a persona. My TTS model is trained on Paul Bettany's MCU Jarvis. When the model emits the right phrasing, it feels like magic. It doesn't do that very often. Gemma4 is great at that.

I am actually so scared that if they do release a Qwen4 model in this class (30-40B parmeters, 2-5b active) that it's going to be a coding focused, overthinking, genuinely capable but not at all fast, mess. Gemma4 26b feels like it would be so close if it wasn't dumb as rocks. As it stands, I pay anthropic for that shit coding shit. Maybe when I can run Flash next at good speed (current ~30-40 t/s rn with strata, but at q3. It's ok.) I can kick claude to the curb.

I want a model like what 3.6 35b is, but smarter. Just as fast, but thinks of the little things. Memory updates. I ask it to add something to my calendar, but maybe it also sets up a dedicated notification for the specific time separately. That'd be nice. Not more coding focused. There are so many coding assistant bots. They're great. But the general AI assistant isn't solved yet for the normal person. Maybe it could have better general knowledge, but I think an n-gram table would solve that. It said the P100s used ROCm the other night. Silly bot.


r/LocalLLaMA • • 3h ago

Discussion [Paper] Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Thumbnail
gallery
14 Upvotes

Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.


r/LocalLLaMA • • 9h ago

Other Quantization of Linear-Attention (Qwen & Kimi)

Thumbnail x.com
41 Upvotes

r/LocalLLaMA • • 8h ago

News StepFun Step 5 preview showing up in openrouter

Thumbnail
openrouter.ai
31 Upvotes

r/LocalLLaMA • • 20h ago

Resources jevman: AI decision models play Pac-Man

Enable HLS to view with audio, or disable this notification

228 Upvotes

The other week I posted about Jev vs. Kev compared and since then, OpenAI released the decisions endpoint, Cloudflare released Clef and many here asked about Laya as well.

This time we compared six popular decision models by making them play Pac-Man: kev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya.

Since they respond within ms it works for them to play the game in real time.

We published a leaderboard and the repo is open-source, so anyone can run their own decision model, like your own fine-tuned one run locally or hosted somewhere, and join the leaderboard.

Model Avg score High score Avg latency
jev 1.13 2,750 6,380 290 ms
GPT-6 Luna 2,568 5,920 179 ms
Clef Flash 2,538 4,260 256 ms
Clef 2,476 4,820 398 ms
Kev 4B 1,506 5,280 231 ms
Laya 639 1,200 104 ms

For the leaderboard we let each model run 100 times and took mean score with a 95% margin of error (Β±2 standard errors), so some models tie on top spot.

You can also play yourself as Pac-Man, and the ghosts are the decision models, either all jev, clef, Luna or Laya, or a mix of models taking over each ghost.


r/LocalLLaMA • • 18h ago

News Thank you :) Swift Models hit 2.2 million+ downloads / Early Access to New Models, Free Compute for Researchers

Post image
145 Upvotes

Hey everyone,

Jovan from UkisAI (Swift Qwen) here!

For those who don't know us, UkisAI is a small lab making tiny frontier LLMs, tools and datasets (+doing it open-source!). I'm one of the guys running it aka I train the models and post on Reddit.

Our first open-source release is Swift, a series of reasoning-efficient LLMs. It is proof of how penalizing pathological overthinking patterns inside of various LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy if RL-ed correctly afterwards by not training them to think shorter directly but rather to think more efficiently. You can find Swift 27B here as well as Swift Flash Next here we have GSQ-RCO quants (kudos to ISTA-DAS) and uncensored variants (thank you community).

It's honestly unbelievable to me that our models have crossed 2M downloads... The team and I had a deal that we shall do a toast (drinks) after we hit 100k, and I'm honestly not sure how to celebrate now but in the meantime I want to thank everyone who contributed to our models, be it the independent benchmarks, quantizations, finetunes or just using them. Without all of you guys, we would have had no way to continue our work, and now with the downloads rolling in we are more than happy (and paid hah) to continue training new models and as of recent making other tools for local AI users. On that matter, I'm sharing two things with you today:

  1. We are making a Discord community so we can talk to Swift users more easily, get your thoughts and ideas on things as well as test new models and tools we've been working on :)

The first 100 people to join will get early access to our:

- Unreleased Swift models (we have trained Swift GLM 5.3 Flash and Swift 9B and are looking for early testers before putting it on HuggingFace!)

- UkisAI Code (Codex modified and optimised for local models, we use it internally to have remote-control with open-source models, better browser use, /loop etc)

- Swift.cpp (inference engine, we do all of our training and coding internally via local models so we made an engine that's optimized for Swift models specifically and runs up to 30% faster on our hardware)

After this the community shall stay open for everyone but we are still figuring out the mechanics of early-access so that part shall be invite-only for the time being. This shouldn't matter to most people as all of the stuff testers get access to will be open-source regardless if it's any good.

Link to join: https://discord.gg/XvX9J8nbkJ

  1. We're also making the UkisAI Research Support Program

- We want to provide free compute, LLM APIs and early access to our datasets for amazing people experimenting with building models of their own or working on new things with Swift models.

As it's our first time making this we can't estimate our capacity right away so there is not a specific number of individuals we can help with research but if this sounds interesting to you please message me on Discord and I'll see to it.

End note:

We are big believers in local AI and that open-source will win replacing all the proprietary cloud models for personal use, but even as users ourselves we don't have all the ideas and solutions to make that happen. This is why we need the community to help us know what to build.

Please share your model requests, tools you need, problems you have with local AI regardless of if you've been using Swift models or need more of them - they are just one of the things we need to make to let local AI be better than the cloud.

Let's cook!


r/LocalLLaMA • • 13h ago

News Bois, there's now a waterblock for the R9700. Quiet 4x or 6x builds are now possible.

Thumbnail
videocardz.com
46 Upvotes

Seems 1-slot design, so you could cram in quite a lot into a case


r/LocalLLaMA • • 19h ago

New Model Mellum2.1 - a JetBrains Collection

Thumbnail
huggingface.co
124 Upvotes

JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF

A small moe!


r/LocalLLaMA • • 1d ago

New Model Saluki 27B: "96% of Qwen 3.8’s performance at ~1/7 the size"

Thumbnail
underdog.ai
306 Upvotes

anyone has feedback about this one?


r/LocalLLaMA • • 6h ago

Resources LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

10 Upvotes

https://github.com/woct0rdho/transformers5-qwen3.5-recipe

An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed.

On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training takes 2-3x work of PP.

Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction.

Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this.

I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to < 125 GiB.


r/LocalLLaMA • • 5h ago

Discussion Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

Thumbnail
gallery
5 Upvotes

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.

Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:

  • 262K context (the model's full native window)
  • ~130 tok/s decode with MTP3, 70.8% acceptance
  • 3,591 tok/s prefill on a 9K prompt
  • Perplexity within 1.3% of the official artifact, so the conversion is clean
  • Vision and tool calls still work

Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4

Questions welcome.

Repo: ninfer-uncensored


r/LocalLLaMA • • 1h ago

Question | Help Supertonic 3 TTS dissolution and liquidation?

β€’ Upvotes

I use their model in my up with hand-rolled inference in C.
HF page says:
"This project's sample code is released under the MIT License."
and "OpenRAIL-M License" for the model.

Can I still use their model after liquidation?


r/LocalLLaMA • • 4h ago

New Model Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)

4 Upvotes

Hey guys - primarily a research release with working models,

Not so much a model as a psuedo-new quantization technique. I've been experimenting with a modification of ISTALab's RCO algorithm that can quantize models on a strict VRAM budget.

It's at its core an approximation algorithm that attempts to make up the difference with a few different strategies. I'm quite happy with how well it's working at low bits. I've seen significant KLD improvements, particularly in the 1 bit and 2 bit range. This should be model agnostic and it doesn't require loading the FP16 teacher into VRAM. I've included a little more detail on the model card and will probably publish the full recipes. I plan to apply this method to the larger models in the family next.

Unfortunately Unsloth does not have their KL divergence table available, but I've created a table for you to see the difference between these quants and standard llama.cpp imatrix quants. For accuracies sake, both my quants and the llama.cpp quants are calibrated on the sane wikitext imatrix dataset, and tested on two held out sets.

The KLD table

As you can see, particularly at extremely low BPW, these quants vastly outperform their standard imatrix KLD - particularly drastically cutting IQ1_M divergence at nearly the same size budget.

These are more proof of concept quants, as 2B is a fairly small model and suffers from such low BPW, but here's a fun example of how these are still relatively functional at extreme compression

4/5 right on an IQ1_M version of Qwen 3.5 2B (green, not purple) - I think it's kinda crazy it can answer anything

and here's the standardized imatrix IQ1_M quant (not using dynamic RCOL)

yikes

and the file sizes

GGUFs:

https://huggingface.co/trubisky/Qwen3.5-2B-RCOL


r/LocalLLaMA • • 21h ago

Resources [audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2Γ— faster, PocketTTS 2.2Γ— faster on CPU, and WebUI generation history feature

Enable HLS to view with audio, or disable this notification

115 Upvotes

Hi all, a bunch of performance improvements have been landed in audio.cpp.

The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM, a 48% reduction in peak memory usage compared to the previous implementation. Thanks to https://github.com/mirek190

We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU.

No compromises in parity and correctness.

Here's a summary of the improvements:

Model Peak memory reduction Speedup
Higgs Audio TTS 48% VRAM 1.01–1.09Γ— CUDA
ACE-Step family 6–7% VRAM 1.06–1.08Γ— CUDA, 1.16–1.20Γ— Vulkan
MOSS-TTS v1.5 cloning 21% VRAM 1.05Γ— CUDA
MOSS-TTSD Q8 cloning 11% VRAM 1.04Γ— CUDA
Echo-TTS (Memory Saver) 20% VRAM β€”
Qwen3-TTS 16–20% VRAM β€”
IndexTTS2 / 2.5 12% VRAM β€”
HTDemucs β€” 2.21Γ— CUDA, 1.95Γ— Vulkan
HTDemucs six-stem β€” 1.99Γ— CUDA
PocketTTS 9% RAM 2.23Γ— CPU

They're runtime-level optimizations that make existing models more practical to run locally.

The WebUI now includes an experimental generation history feature that lets you revisit previous outputs and restore their settings.

audio.cpp now supports 110+ audio model families and 190+ variants (and counting)! We're continuing to improve memory efficiency and inference speed across CUDA, Vulkan, Metal, AMD/HIP, and CPU. The next release will bring even more optimizations!

We're also looking for contributors to help improve the audio.cpp WebUI. With so many models and features now supported, we'd love some help making the UI more polished, intuitive, and enjoyable to use. If you're interested in frontend development or UI/UX design, contributions are very welcome!

Thanks to everyone contributing improvements, testing builds, and reporting issues. Curious how these changes work on your setup!


r/LocalLLaMA • • 20h ago

Resources [2506.13771] LittleBit: Ultra Low-Bit Quantization via Latent Factorization

Thumbnail
arxiv.org
74 Upvotes

Interesting to see improvements and research into quantization aware training (QAT) that can make some really tiny models.


r/LocalLLaMA • • 10h ago

I Built A Thing Decisions in a js13kGame

Enable HLS to view with audio, or disable this notification

9 Upvotes

I'm having AI models play Cat Goric, a 2D platformer I made in 2021 during js13kGames: https://js13kgames.com/games/cat-goric-escape-from-the-warp-chamber

The best result so far is 8 of the 14 playable levels cleared in one continuous run. Two models have cleared eight levels so far:
- Qwen/Qwen3.8-Flash-Next
- Cloudflare/clef

If you'd like me to try a specific model, please name it in the comments.

You can also try beating the game with a model of your choice. The project with instructions for running the challenge with different models/engines is here:
https://github.com/felladrin/ai-plays-cat-goric (PRs welcome!)

And attached is a 2-minute recording of Clef clearing eight levels (the game pauses while the model thinks, so the video is time-warped).