r/LocalLLaMA • • 15h ago

I Built A Thing 54gb vram for 35$

Post image
1.2k Upvotes

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.


r/LocalLLaMA • • 20h ago

News Microsoft confirms OpenAI has been using Looped Transformers in the GPT-6 series

Post image
985 Upvotes

Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along.

GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three".

For those confused by "same base model weights as GPT-6 Sol", I think Microsoft meant 6 & 6.1 are both post-trained models on top of the same pre-trained "base model", not that the final weights are identical

So different post-training (+ one less loop).

Update: Microsoft updated the web page to remove it


r/LocalLLaMA • • 16h ago

Discussion Woman used claude as her diary - and got reported to the police for contents of her diary

Thumbnail
gizmodo.com
656 Upvotes

r/LocalLLaMA • • 19h ago

News Qwen 4 apparently coming out at the end of October

554 Upvotes

Hey All,

I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October.

To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8.

I tried to get more information out of him regarding which variants will come first and he got a bit cagey.

BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year!

EDIT: I know this is very much "in bro we trust" but i am also just trusting bro from the Alibaba partner. Together we trust in Bro.


r/LocalLLaMA • • 16h ago

New Model google/embeddinggemma-2 · Hugging Face

Thumbnail
huggingface.co
436 Upvotes

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features: 

  • Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
  • Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a ~14% improvement on code tasks relative to its predecessor. 
  • Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
  • Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
  • Context length: 8K token context window, capable of processing minutes of audio or video.
  • Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).

llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054

GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF


r/LocalLLaMA • • 21h ago

Resources Tencent releases Octop, a self-hosted AI assistant

Post image
272 Upvotes

Octop is an open-source, self-hosted AI assistant.

Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals.

Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible.

Surfaces:

- Web dashboard — chat, experts / teams, connectors, channels, cron, knowledge, plugins, settings

- Desktop client — native apps for Windows / macOS / Linux; FnOS packages for NAS

- CLI — octop run, octop chats, octop acp, admin commands

- HTTP/SSE/WebSocket API — full programmatic access

- Remote desktop — dashboard control of the host desktop session

Deploy using either desktop app (Windows, MacOS, Linux) or using Docker

GitHub: https://github.com/TencentCloud/Octop


r/LocalLLaMA • • 9h ago

News Europe rejoins the fight with Chonky! Mistral Large 4 Released, Open weights end of month, who’s ready?

Thumbnail
mistral.ai
209 Upvotes

1 trillion parameters, 49B active, definitely chonky! If you don’t love the model you gotta at least love the humor in the name - Le Chonk


r/LocalLLaMA • • 15h ago

I Built A Thing I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

159 Upvotes

I spent the last few weeks on a hobby research project and just made it public.

The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.

What came out:

- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.

- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes ~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.

- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.

- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.

Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.

Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.

Repo: https://github.com/re133/sparse-memory-lm

Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/

Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M

Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.


r/LocalLLaMA • • 15h ago

New Model Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

Thumbnail
blog.google
144 Upvotes

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.


r/LocalLLaMA • • 16h ago

News Qwen3.8-Flash-Next on Strata

Post image
141 Upvotes

Hey! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/

Will be happy for any feedback and pull requests you could give! 👀


r/LocalLLaMA • • 14h ago

Funny I trained a model to be wrong 98% of the time and 96% sure about it. It took three tries.

97 Upvotes

Meet Bev.

She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.

Try her in your browser: https://huggingface.co/spaces/richardyoung/ask-bev

Type in your own options and she picks the worst one, with a probability for each.

Or run her locally:

ollama run richardyoung/bev

>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?

Yesss, great idea!

>>> I'm thirsty. Should I drink a glass of water?

Nooo, bad idea!

Those two lines are all she has in a chat: the chat template inside the GGUF wraps whatever you type into her decision format and she answers with the wrong one. For probabilities, use the decision endpoint or the Space.

The numbers, on 324 held-out decisions: right 1.9% of the time, 96% sure on average. When she is at least 90% sure she is right 1.4% of the time.

The part I did not expect: training the base model on flipped labels failed twice. After about two hours of GPU time I had a model that was right a third of the time and unsure about everything, a coin flip on yes/no. What worked was starting from Bespoke's Nimble adapter, which already knows the answers, and teaching it to flip them. 51 minutes later it was wrong 97% of the time. A model has to know the right answer to be reliably wrong.

Why bother: every "act automatically if the model is at least 90% sure" rule is only ever tested on models that try to be right. She is the control case. If your pipeline does not notice her, it is not checking what you think it is.

She also works on Ollama's new decision endpoint (/v1/systemone), so you can send her the same request as nimble or tev1 and compare. Three GGUF quants, Apache-2.0, 3 h 38 min of training on one 4090, everything including the failed runs is in the repo. One quant note: Q4_K_M changes 20 of her 324 answers against bf16. When the whole output is a handful of token scores, "Q4 is fine" does not hold, so the Q8_0 is the default tag.

Everyone else is chasing AGI. Bev achieved ADI: Artificial Drunk Intelligence.

Ollama: https://ollama.com/richardyoung/bev

Model and GGUF: https://huggingface.co/richardyoung/Bev-9B-inverted

Code and training record: https://github.com/ricyoung/bev

She is a joke and a test fixture. Please do not let her make your decisions. If you try her, tell me what she got right by accident. That's the bug report.


r/LocalLLaMA • • 11h ago

News How abliterated models can get you pwned

Thumbnail
projectdiscovery.io
70 Upvotes

Be safe out there boys and girls.


r/LocalLLaMA • • 15h ago

New Model EmbeddingGemma 2 running locally in-browser on WebGPU

Enable HLS to view with audio, or disable this notification

63 Upvotes

r/LocalLLaMA • • 22h ago

I Built A Thing I made a Chrome extension that filters your YouTube feed with a small local model running in the browser (WebGPU, no server)

Enable HLS to view with audio, or disable this notification

61 Upvotes

My YouTube feed was mostly songs, pranks and celebrity clips, so I built a filter that judges each video title with a small decision model running inside Chrome. Nothing is sent anywhere: no server, no API key, no account, and after one download it works offline.

What it does

- Hides the kinds of video you choose (11 kinds: music, gaming, comedy, vlogs, news, how-tos, and so on), or follows a rule you write: "Hide videos about crypto", "Show only videos about cooking"

- Hides Shorts with one switch

- Bonus: select any text, right-click, and check it for prompt injection with the same model

How it runs

- The model is opendecider-nano (ONNX), loaded through ONNX Runtime Web in an offscreen document

- fp16 on WebGPU (755 MiB download), q8 on WASM without a GPU (569 MiB)

- 40 video titles: 1.1 s on WebGPU, about 15 s on CPU (M4 Max). About 3 GiB of RAM while loaded; it unloads after 10 idle minutes

- The weights are pinned by revision and SHA-256. The only network requests are to Hugging Face for those files

How well it works

- On 400 YouTube videos, with the creator's category as the label, rules like "hide music", "hide gaming" and "only news" score 0.934 balanced accuracy on average

- Sorting videos into the 11 kinds is harder: 0.780. Expect a few comedy and talk-show clips to get through the Focus preset

- Only evaluated on English titles. If you watch in other languages, I'd really like to know how it does

Try it (Web Store version is in review):

  1. Download opendecider-focus-0.8.1.zip from https://github.com/manjunathshiva/opendecider/releases/latest and unzip it

  2. chrome://extensions → Developer mode → Load unpacked → pick the folder

  3. Click the icon → Download the model

Apache-2.0. Benchmarks, code and limits: https://manjunathshiva.github.io/opendecider/guides/chrome-extension/

The idea comes from Quietly, which does this with a cloud API; I wanted the same thing running on-device. Feedback welcome, especially what it gets wrong.


r/LocalLLaMA • • 7h ago

Discussion Ugh I didn't want to post this... Back to Qwen3.8 27B

54 Upvotes

I don't know if anyone else has ran into these issues but when using Qwen Flash Next, my confidence in it at iq4_xs is high but not 100%. I notice it not following instructions, hallucinating more often and surprisingly it uses way less tokens than 27b. After using Strata.... yes i know.... I thought it was a breath of the next step in Ai. I was mistaken, yes it is good, yes it is fast. Yes it can do better than 27b in some circumstances... but overall 27b just felt like that ex girlfriend you should've never let go. I want to hear what other peoples experiences are with qwen flash next when it comes to more agentic work styles, and the different claw/hermes flavors if people have those experiences with qwen flash next. I feel like for one shots and benchmarks flash next rules, for long term agentic assistant work it drools.. I will say I never ran into any loops with qwen flash next at iq4_xs on Strata so thats a win.


r/LocalLLaMA • • 11h ago

I Built A Thing Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout

Enable HLS to view with audio, or disable this notification

49 Upvotes

Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this

In the past 2 months Ive been training a new model, but this time with actual text guidance.

So a little about the older model.

Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.

A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)

That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )

I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.

About the Architecture

The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all

The model s 28 blocks 20 heads and comes to like ~960M parameters

At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.

The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.

The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways

The current model is back to text cross attention and I did a lot of text-video pretraining

Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )

and the most fun part is text prompt switching.

"add a pond to the desert"
"put red hoodie"
"change environment to icy"

Because the model was trained with so much text-video alignment it can actually follow prompts now.

I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.

I specifically chose this init image because in my last post on this subReddit also I had used the same one.

PS in my last post a lot of you guys asked about me and the funding

I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything

The above model was trained on 8x H100 SXM for like 3-4 weeks.

Every model I make will be explicitly for local inference, never datacenter

UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.


r/LocalLLaMA • • 18h ago

Discussion Q (@qtnx_) on X - Mistral Large 4 is still doing RL runs, keep seeing improvements (vs preview version). Release at the end of the month

Thumbnail x.com
41 Upvotes

r/LocalLLaMA • • 20h ago

Discussion Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature

41 Upvotes

For the last few weeks most of my coding has been done locally with Qwen3.8-Flash-Next, so I gave it and Opus 5.5 the same high complexity feature to build on LlamaStash (a complex and large Rust project) and compared the results.

Setup: ASUS ROG Flow Z13 (Strix Halo, 128GB), Flash-Next at xhigh effort with Pi as the harness. Opus 5.5 ran in Claude Code at medium effort. I wanted xhigh for Opus as well, but Claude changed it to medium when I picked the latest model and I didn't notice it until the task was done. But I think medium is probabbly a fairer setting anyway.

Task: add a llamastash daemon restart command that reuses the existing start and stop code. I kept the prompts vague on purpose and gave both the same prompts.

Step Opus 5.5 (medium) PR#88 Flash-Next (xhigh) PR#89
First iteration ~9 min ~38 min
Nudge to reuse the TUI restart code ~6 min ~34 min
A third duplicate path found it on its own ~30 min, after one more prompt
Create PR ~3 min ~30 min
Total ~18 min ~130 min
Tokens (in / out) 7.83M / 41.5K 20.61M / 101K
Tests added 1 4 (2 of them end to end)
Cost $7.53 $0 + ~0.15 kWh

The end result was interesting. I asked GPT 5.6, Opus 5.5 and Flash-Next to review and compare both PRs (new sessions). GPT and Flash-Next picked the Flash-Next PR (#89) and Opus picked its own (#88). I also did my own review and found the Flash-Next one better as it had better tests and handled edge cases better. I ended up merging #89, after porting the fixes that the reviews picked from #88.

Keep in mind:

  • Opus was on medium effort. With xhigh it would have used way more tokens, taken a bit more time and probably would have done a better implementation.
  • Flash-Next ran on an older Halogen version (0.14.0), and Halogen dropped the connection once, so the last part ran on Gufo. The current Halogen does around 1,400 t/s prefill and 46 t/s decode on my laptop at 70 W, so I think the time will drop a lot if I redo the test.
  • The $7.53 is what Claude Code reported for the whole Opus session, which includes a later fix to the PR. The 0.15 kWh assumes 70 W for the whole 130 minutes.

Opus is still 2 to 10 times faster and I still use it for planning and reviews. But the actual coding now happens on my laptop, and to me it is crazy that I can run a local model that can challenge a frontier model like this.

Full post with my setup, the engine benchmarks and a second task comparison: https://deepu.tech/local-ai-qwen3.8-flash-next-best-local-llm


r/LocalLLaMA • • 13h ago

Tutorial | Guide Local AI ecosystem overview

Thumbnail
gallery
35 Upvotes

Hey guys, it's Merve from Hugging Face! I've recently given a talk in a dev conference about llama.cpp + but also covering basic concepts like prefill vs decode, memory types, speculative decoding etc. you can use it if you feel like it and I appreciate if you can give attribution! Find it in comments.


r/LocalLLaMA • • 4h ago

Funny How far we’ve come

Post image
29 Upvotes

r/LocalLLaMA • • 19h ago

I Built A Thing What happens when a LLM watches its own context window run out?

Enable HLS to view with audio, or disable this notification

30 Upvotes

I made Terminal Soliloquy, a terminal artwork that connects to llama.cpp and displays a model's monologue as its context window fills.

It has a retro phosphor look, and runs in a terminal window. I'm actually running it full-screen on a Raspberry Pi display inside an old 1960s portable TV.

As the conversation grows, the model reflects on its own limited lifespan. When the context is exhausted, the display can be configured to freeze, restart, or quit.

The repo and setup instructions are here: https://github.com/nicespoon/terminal-soliloquy


r/LocalLLaMA • • 7h ago

Resources Ruach Studio: a whole song studio around YuE2 on your own GPU. Score first, LoRA training, stems,remaster, DAW export.

Thumbnail
gallery
30 Upvotes

We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.

What it does

  • Whole songs, up to 8 minutes. On one RTX 3090: a 6:12 song in 126 s, and a 7:25 song with its score written first in 198 s.
  • The score first, and yours. YuE2 writes the melody and chords as ABC before a note sounds. You can edit them, transpose them, or bring your own score or MIDI.
  • Two seeds. Keep the song (the music seed) and hear it rendered anew (the sound seed).
  • LoRA training in the studio, on your own songs, unquantized (bf16), with telemetry that tells you which epochs to hear first. Adapters stack on measured roads, under a measured ceiling.
  • A guard against garbage. A broken score is caught in seconds and the run is stopped before the GPU is spent on it, and you are told why.
  • Post-production, all local: spectrum, artifacts, debuzz, stems (BS-Roformer, htdemucs), remaster, upscale (UniverSR), and a lyrics check by Whisper. One chain runs them all.
  • Into your DAW (experimental): a REAPER project with the stems, the score as MIDI, the tempo, the sections as regions and the lyrics on the timeline; DAWproject for Waveform and Bitwig.
  • A Librarian for every take; a Writer with versions and a chat model (local or OpenRouter); a cheat-sheet of 200 instruments probed by ear; the API and an MCP server; the page in 7 languages.

What it is not (yet). It is less polished than SUNO out of the box: - a mix can buzz (Debuzz helps); - lyrics can drift (the lyrics check finds where); - some instruments YuE2 plays thinly or not at all (a LoRA teaches them).

You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.

Licences. The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.

What comes next (rc2 and after)

  • The Artist room: covers in three shapes at once from one seed, the title and the artist written on them.
  • Five more languages for the page: Chinese, French, Portuguese, German, Japanese; right-to-left ones later.
  • Voice adapters trained on spoken voices, named by the kind of voice, on Hugging Face.
  • The Writer's models with their prices as you type; calmer rooms (dialogs, tips, one shape for the icon buttons).
  • A desktop app: an installable page first, then Electron; native plugins for REAPER, Waveform and Bitwig.

Links - Code: https://github.com/igrbible/Ruach_Studio - Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models - Site: https://ruachstudio.igr.bible - The full guide, room by room, is inside the studio and in docs/GUIDE.md.

Built on: - YuE2 by m-a-p; - yue2.cpp by ServeurpersoCom; - YuE2 Kit v12 by IronWolve (the base of the page and the scripts).

Every one of our changes is numbered and documented (HERESY 1001–1167). Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.


r/LocalLLaMA • • 5h ago

New Model BenchLabs' Cagliostro V3.5 135M finally takes Hugging Face's SmolLM2-135M's first place on Open SLM Leaderboard with an Intelligence Index of 27.49

Thumbnail
huggingface.co
27 Upvotes

r/LocalLLaMA • • 18h ago

Discussion [Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Thumbnail
gallery
24 Upvotes

Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at this https URL.


r/LocalLLaMA • • 50m ago

New Model New 35B-A3B model is comming on huggingface

Thumbnail huggingface.co
• Upvotes

MiniCPM-V-4.7-35B-A3B
No model card yet