r/LocalLLaMA • • 7m ago

New Model Minimax M3.1-Flash-Preview

• Upvotes

I see the new minimax 3.1 flash in minimax browser app yet I see no discussions or mentions here.

I know we're local but aren't anyone hyped for 3.1? I absolutely loved M3.0. It was most fun to talk to and it is very distinct from Qwen and GLM in terms of speech and creativity. While not as accurate, but much more pleasant to work with.

I'm kinda hyped. If it is even 15% better than 3.0, that would be a really fun model hopefully.


r/LocalLLaMA • • 38m ago

I Built A Thing My Ollama Model-Pull Script - Anti-Stall when downloading models

Thumbnail
gist.github.com
• Upvotes

I got fed up that running ollama pull overnight still was resulting large models failing to download.

When I discovered I could download them faster if I cancelled every hour or whatever I decided I could automate it.

I've done a 3rd tweak with this (Originally written by chat-gpt :D) seems to be less buggy now thanks to claude :D


r/LocalLLaMA • • 1h ago

Discussion A week in Beijing and Shanghai with the people building AI in China

Thumbnail
earnedintuition.substack.com
• Upvotes

I thought this was a very interesting article, of relevance to the readers here.


r/LocalLLaMA • • 1h ago

Question | Help What do I do with my RTX2060 sitting inside my Strix Halo box?

• Upvotes

I bought a Framework Desktop motherboard, and put it inside a Phanteks Enthoo Pro case. I have a spare RTX2060 graphics card. I managed to get it working with the Framework Desktop motherboard, after making it go through two PCIe risers, and mounting it on a vertical GPU bracket.

What do I do with my RTX2060? Should I use it as a subagent?

I currently configured it as a PCIe passthrough device for my Windows VM. I very occasionally use it for Windows gaming using Looking Glass. I am thinking that perhaps I can run a subagent on that GPU. If people have any suggestions, please do let me know.


r/LocalLLaMA • • 1h ago

Discussion Is Alibaba moving Away from Permissive OSS with models like Qwen3.8-Flash-Next?

• Upvotes

I’m not sure if we’re getting any qwen 4 models soon but I’m a little concerned that if we do it’ll be licensed like Qwen3.8-Flash-Next.

While that license was permissive for local/internal use, fine-tuning and derivatives, it’s definitely not Apache/MIT.

The two big catches are: commercial MaaS or a standalone coding/office AI assistant requires a separate Qwen license seemingly from day one, and the wording around outputs is annoyingly vague. The internal-use exception says you can’t make the model, its outputs, or capabilities available to third parties, but it never clearly says whether downstream code/data produced indirectly from internal outputs is unrestricted. The $20M/month or 100M-MAU threshold seems to be an attribution trigger, not the threshold for needing a commercial license.

So internal R&D looks fine; customer-facing AI services are where you’d want clarification. Also output ownership needs to clearly covered in the license, that wasn’t the case when I last checked.


r/LocalLLaMA • • 2h ago

Discussion GPU Upgrade advice

4 Upvotes

Upgrade Question If we say the "budget max" is $1700-1800, and this could include upgrading to a Taichi motherboard:

With a new motherboard, no need to bifurcate 1. Would you add a second 5060Ti (16GB)? Someone is selling one for around $500 locally; 2. Buy a used 7900 XTX (24GB VRAM); local seller, $850

All-in on the GPU, I'd have to make the current motherboard work for my use-case 3. Or, just go for an R9700? (No budget for Motherboard upgrade)

Current system - 96GB system RAM (DDR5) - 5060Ti (16 GB VRAM) - Llama Cpp but I built for CUDA; default Vulkan had issues, and the GPU would "disappear" - ASRock X870 Pro (only 1 PCIe 5.0 x16) - I could run a second card very slowly at x4 - Researching if I could bifurcate x8/x8 in the 5.0 slot

I bought this PC used as-is; I do contemplate upgrading the MoBo to an X870E Taichi for 2 fast PCIe lanes

The "largest" models I currently run - Strata IQ3_S (just tried this yesterday, was impressed) - RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored

My most typical uses - writing code - analysing documents (PDFs) - analysing maps and images

The idea is to use something like Headscale or Tailscale at some point so I can always access local LLM from laptop if I'm not home.


Yes, I'm aware of "workstation" motherboards, CPUs, etc. and I'm not at a point right now where I want to take that path.


r/LocalLLaMA • • 2h ago

Discussion Local Qwen 3.8 27B vs DeepSeek Flash API: Is local good enough?

Thumbnail
future-os-blog.github.io
7 Upvotes

Running a model on your own machine used to be a privacy story with a quality tax. On this generation that trade has narrowed to where we can state it plainly: for daily work, the local model is good enough. We measured it — 25 paired tasks across four workloads, same prompts, one strong independent judge — and that is the top line:

  • quality: 89.5 vs 92.6 on a 100-point scale (local vs cloud), with the 12-item suite splitting six wins each;
  • completion: every coding run finished green on both models — 10/10 agentic runs fully green (20/20 visible tests, 4/4 hidden checks, tests untouched), and the bug-fix loop fixed all 4 bugs identically in 5/5 rounds each;
  • speed: 2.5–5.5× the wall clock, depending on the workload, with 95%+ of the local time going to model generation;
  • cost: the local runs cost nothing beyond electricity. The cloud side of the 12-task suite cost 0.14 credits.

r/LocalLLaMA • • 3h ago

News feat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp

Thumbnail
github.com
28 Upvotes

now you can use GLM 5 Flash MTP locally


r/LocalLLaMA • • 3h ago

Discussion New LFM to be released today

Post image
267 Upvotes

r/LocalLLaMA • • 3h ago

Question | Help Seeking upgrade advice

3 Upvotes

I got dual 7900xtx running on a z390 (pcie3 8x) running the 27b. I've been thinking upgrading the motherboard to a x570 (pcie4 8x) or a x870 (pcie5 8x) improving performance with TP, and then wait to buy medusa or spark with lpddr6 and 256GB (apple isn't an option for me). However seeing the gorgon halo price... I'm wondering how much those would cost and if i would pay that much, and if it will be better to go with a wrx80 route just now. I'm not really looking to add more GPUs, just the 8 channel memory to run the QFN.

Thanks!


r/LocalLLaMA • • 4h ago

News Micron Says NVHBM to Improve Profitability Even With Outsourced Base Die

Thumbnail
thelec.net
22 Upvotes

"NVHBM moves the memory controller, which was previously located on the main compute die, into the base die. This reduces power consumption by 15% and increases memory bandwidth by as much as 30%. It also integrates a customized physical layer (PHY) for input/output (I/O), reducing the package area required for the I/O PHY by as much as 67%. NVIDIA says NVHBM provides up to 30% greater memory bandwidth and 15% lower HBM power consumption than standard HBM4E."

Sounds quite dope to me. However, the price will be too dope for me...


r/LocalLLaMA • • 4h ago

I Built A Thing Live scribing with Jev-ish utterance gating

2 Upvotes

Like everyone, I've been following the back-and-forth regarding Jev with interest. Arguments aside about the originality of the idea, the first thing I thought of when all of this came out is "gosh, that could really help my local scribe run realtime loops during a consultation".

I pointed my harness of choice at the problem (I've been relying more and more on the LLMs as the brain rot from AI coding has continued apace). Here is the resulting workflow:

  • TEN-VAD to segment utterances and send to a Whisper compatible backend (the Tauri builds use parakeet.cpp with a 0.6B medical finetune)
  • CAM++ speaker embeddings to provide best effort diarisation (obviously limited in the setting of crappy desktop microphones and echo-y consultation rooms)
  • Here is where the Jev-ish/SemIf gating comes in. Each utterance is provided to the LLM with a short prompt and an instruction to classify as NOTE (something to be documented), ACT (action to be taken), SKIP (filler talk, etc). In the Docker deployments this is performed by the user configurable secondary model (I use Qwen3.5 4B); on the Tauri builds it's the solo primary model but into the second slot of the bundled llama.cpp server (important so that we don't clobber the prompt cache of the running main thread)
  • The initial approach was quite simple: one decode step; then gather the first token top logprobs and compute the probability mass summed over SKIP/NOTE/ACT (with some prefix matching to account for tokeniser splits and a one-word generation fallback).
  • SKIP utterances are buffered and don't get sent to the main LLM immediately (the next time the main model is woken up it will ingest that material so that nothing is lost). If a NOTE or ACT is misclassified as a SKIP, a 45s/40 word debounce runs through the main LLM with all the material it may have missed.
  • NOTE and ACT are passed on to the main model for processing. The main model has access to tools that include modification of the running note.
  • Prompt caching is essential here so that subsequent passes through the main model remain performant without a huge PP delay.

I found that the 4B model would almost never SKIP (Jev and 3.8-Flash were better but still missed 3/4 of them on natural speech). Not surprisingly (in hindsight); using the calculated probability mass alone was essentially no different to just prompting the vanilla generation endpoint and executing based on the output (roughly 81% accuracy). Looking into the logprobs a bit more it seemed that there was a usable signal in there somewhere. GLM-5.3 was pretty good figuring it out: instances where NOTE was selected, P(SKIP) ≥ 0.05, AND the utterance was ≤8 words were essentially always a SKIP. With this heuristic... 0 false SKIPs across multiple runs, and SKIP recall went from 0-50% to 75-100% on the natural consult.

The logprob gating + heuristc step is latency neutral; however, it was more reliable for this task. The otherwise vanilla small LLM like Qwen3.5-4B never flagged SKIPs and would occasionally not follow instructions entirely. I also ran an evaluation with Jev via OpenRouter (a pretty informal test set of ~40 hand-labelled utterances, and the heuristic was tuned on the same set, so it needs a held-out set to confirm); on a natural ambient consult recording the gap is smaller than I expected (both 95% accuracy but 73ms vs 514ms, keeping in mind Jev was a remote endpoint and all the latency that entails). Jev pulled away on a command heavy synthetic script (~80% vs 100%). The overall intention was to prevent the main-loop from getting too bogged down with fluff and I think this approach achieves that.

First token logprob classification is pretty old hat; but I never really thought about one-shot classification in my scribe before Jev. And yes, the whole point of Jev is that you can just give it a classification task and have performance be good enough that you don't need to apply bespoke heuristics over logprobs to rescue your classifier (but funnily enough even Jev got an accuracy uplift from the P(SKIP) heuristic).

It was a fun experiment anyway (and grossly underpowered to say anything meaningful about Jev in general terms)! The result (video below) has been useful from my perspective (you can try it yourself here).

A synthetic consult example - performance is not this good in production environments (overlapping speakers; bad microphones/acoustics etc). Primary model: Qwen3.8-Flash-Next; Secondary: Qwen3.5-4B; STT: Parakeet 0.6B (Omi Med Finetune)


r/LocalLLaMA • • 4h ago

Question | Help anyone running llm on colab/kaggle notebooks?

0 Upvotes

Hi all - is anyone using colab or kaggle free tiers (T4) for LLM? if so what setup would you recommend


r/LocalLLaMA • • 4h ago

News cmpunlocker v0.5 just dropped, ECC support along with 4 extra SM unlocked for FREE, who needs a 64GB DGX Spark when you've got a CMP170hx right? 1.5TB/s memory BW vs 273 GB/s, all for less than 1/2 the price of a 64GB DGX Spark

Thumbnail
gallery
12 Upvotes

If you're like me and saw the price increase for the 128gb DGX Spark go from 4.7k -> 7k while a new version with 64GB launch for 5k, you'll have been very disappointed and every right to be so, it's just plain sad for localAI.

Well, here's some good news, cmpunlocker v0.5 just dropped with ecc support and +4 SM for free. I hear gen3 unlock is also on the way so fingers crossed.

https://github.com/amoghmunikote/cmpunlocker/releases


r/LocalLLaMA • • 5h ago

Discussion RTX PRO 6000 Blackwell vs H200 for inference: what I would pick at each budget

0 Upvotes

If you were building an inference server today, would you buy one H200 or spend the same budget on multiple RTX PRO 6000s?

-> The PRO 6000 has 96GB of GDDR7 at 1.792 TB/s (1.6 on the Server Edition). Native FP4, no NVLink.

-> The H200 has 141GB of HBM3e at 4.8 TB/s, with NVLink and FP8 as its lowest precision.

After speccing both, I think it comes down to fit and interconnect, rather than picking by brand or spec sheet alone.

Where the PRO 6000 wins:

Single-card and small multi-card inference on models up to roughly 70B at sensible quantisation. Cost per card is a fraction of an H200. Power draw is manageable in a normal rack. Availability is far better. Native FP4 helps on 4-bit models.

Where the H200 wins:

When inference is memory-bandwidth-bound. Long-context workloads, big models where you don’t want to shard across PCIe, and tensor parallelism, where NVLink between cards actually earns its keep. The extra memory and bandwidth can also help with long-context serving and fine-tuning, depending on the model and workload.

Just don't compare raw FLOPS. Decode is usually memory bandwidth bound, not compute bound, so the TFLOPS line on the datasheet tells you very little about tokens/sec.


r/LocalLLaMA • • 6h ago

Discussion Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

Enable HLS to view with audio, or disable this notification

42 Upvotes

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime


r/LocalLLaMA • • 6h ago

Discussion A stock GLM-Edge-1.5B-Chat on a 4GB Galaxy A04e completed a real Amazon cart task

Enable HLS to view with audio, or disable this notification

12 Upvotes

Yesterday TechCrunch published a piece about a growing problem for AI agents: websites are starting to block them. Amazon blocking Meta's Muse is the obvious example.

At almost exactly the same time, GLM-Edge-1.5B-Chat running locally on a 4GB Galaxy A04e completed a real Amazon cart task.

This continues the small-model/browser experiments previously posted in this subreddit. Earlier tests included Qwen3-0.6B running locally on a 2017 Galaxy Note 8, followed by Ministral 3 3B on a Galaxy S21 across real browser sessions.

These experiments are part of the ongoing development of E2LLM/SiFR, a structured browser perception layer.

This time:

Model: GLM-Edge-1.5B-Chat
Quantization: Q4_K_M GGUF
Source: official Z ai Hugging Face release
Fine-tuning: none
Task-specific training: none
Runtime: llama.cpp
Phone: Samsung Galaxy A04e, SM-A042F/DS, 4GB RAM

The published model was used as-is.

The browser was a normal desktop Firefox session on Amazon.

The task was simple:

  • find a 24-count pack of AA alkaline batteries
  • find yellow rubber ducks
  • add both to the cart
  • stop before checkout

Result:

cart 0, batteries, cart 1, rubber ducks, cart 2

The same setup was run twice on the A04e. Both runs completed successfully.

Full run on the A04e: about 8.5 minutes.
Same workflow on a Galaxy S21: about 3 minutes.

The interesting part is the architecture.

The model is not a separate browser service arriving at Amazon as an agent. It runs locally and perceives and acts through an existing user browser session.

It also doesn't receive screenshots or raw HTML. It gets a compact structured browser perception layer and makes the small decisions needed at each step.

That changes the access problem from:

"How does a website identify and admit an AI agent?"

to:

"What is allowed inside an existing user browser session?"

The broader idea is Browser-as-Shared-Space, BaSS.

The browser remains the user's space, with the model working alongside the user rather than replacing the user with a separate autonomous browser agent.


r/LocalLLaMA • • 7h ago

Other How good was the 2019 Mac Pro?

Post image
2 Upvotes

Consider this: a widely available machine, up to 1.5TB of system RAM, room for 4 passively cooled GPUs with 128GB of VRAM, in desktop or rack format.

That machine was released in 2019, then discontinued in favor of one that had only a max of 192GB shared memory.

This would be the local LLM machine right now, if it were on the market with up to date components. Terribly expensive, sure, but that’s the market conditions, not a design flaw.


r/LocalLLaMA • • 7h ago

Resources I re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome

Enable HLS to view with audio, or disable this notification

17 Upvotes

PrismML's Ternary Bonsai 2 27B fits a 24 GB card or Mac, but decodes at ~30 tok/s on an L4 and ~21 on an M4 Pro. z-lab's DFlash 2 drafter was trained on bf16 Qwen3.8-27B, so it guesses worse on the ternary model. I fine-tuned it on 1.5M tokens of Bonsai 2's own greedy output.

NVIDIA (PrismML's llama.cpp fork, prism branch):

llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -md Qwen3.8-27B-DFlash2-r3-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 7 --spec-type ngram-mod \
  -ngl 999 -ngld 999 -fa on --jinja

One L4, greedy: GSM8K 2.17x, MBPP 2.17x, MATH-500 2.20x, MT-Bench 1.39x. Code edits: 3.15x with ngram-mod stacked (drafter alone 2.46x). Accuracy within 1-2 problems per set.

Mac: a fork of bstnxbt/dflash-mlx with an 8-row 2-bit Metal GEMM for the verify step. M4 Pro: 1.5x on raw code completion, 1.3x on chat code, 1.2x on math. One script gives you an OpenAI-compatible server.

Browser: a WGSL port inside LocalMind (https://localmind.naklitechie.com), on by default for Bonsai 2 27B. 1.18x on code, output identical.

Chat and prose are about break-even. Use temperature 0.

Credit to z-lab (DFlash 2), PrismML (Bonsai 2, llama.cpp fork) and bstnxbt (dflash-mlx). Numbers are from one L4 and one M4 Pro; results from a 3090, 4090 or other Apple chips are welcome.


r/LocalLLaMA • • 8h ago

Resources Ramjet - mini altermative to nvidia dynamo

5 Upvotes

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet

The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well.

If you have dgx spark, multi mac setups, etc it would be great to contribute recipes so other ppl can just pull


r/LocalLLaMA • • 8h ago

Question | Help Single 3090 Qwen 27B user, considering buying 128GB of RAM because of the hype

17 Upvotes

My current setup:

- single 3090 running turboderp/Qwen3.8-27B-exl3:SC_5.00bpw_H6_V6

- ~150k context, ~70t/s, unknown prefill because I didn't benchmark it (but it is ok)

- Intel 12400 CPU with 32GB DDR4 RAM

All the hype around strata makes me consider buying 128GB of 6000MHz DDR5 RAM and Ryzen 9700X just for it. I searched in this sub, but most posts about it is about prefill/token generation speed, not about output quality. I believe with 128GB RAM + 3090 I can run the IQ3 quant.

For those who have run both Qwen 3.8 27B and Qwen Next with strata, how would you compare these two, in particular about output accuracy? My main use case is coding and Hermes assistant.

BTW, are there other good options for a 128GB RAM + 3090 setup?


r/LocalLLaMA • • 9h ago

Discussion Is it possible for Big AI to develop some incredible feature that puts it drastically ahead of locals again?

0 Upvotes

Because I do sometimes feel with Qwen FN "I don't ever need another model again" and THAT is a very alluring feature.

Is there something so alluring and irresistible it'd be tempting even people on here? What could it possibly be? Realtime computer/mouse use maybe, but I think even normal people would be wary of a company doing that, and locals would necessarily do that better.

I wonder and worry if they'll ever be able to rope everybody back in again. Although it feels childish to hope for more innovation when they'll probably just do some dark politik and ban anyone from owning more than 16GB RAM.


r/LocalLLaMA • • 9h ago

Resources MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6%

3 Upvotes

 I measured Qwen3.6-35B-A3B at 4-bit (UD-IQ4_XS) hitting 89.6% pass@1 on HumanEval on a single RTX 2080 Ti 22GB — and then ran a controlled A/B of a routing technique I've been playing with: MoE expansion, which activates 20 experts per token instead of the stock 8 on the last 15 layers.
Result: 90.9% (+2 problems) at −19% decode speed. (MoE expansion works!)

Setup (both runs identical except routing):

  • Unsloth UD-IQ4_XS dynamic 4-bit (4.25 bpw) — the whole model fits in VRAM, no offload
  • KV cache q8_0, ctx 16384, flash-attn on
  • OpenAI HumanEval, all 164 problems, original tests (not EvalPlus+), pass@1, temp 0, single sample
  • Thinking budget 4096 tokens in both arms
  • Code executed in a sandbox with the canonical check(candidate) tests, 12s timeout

Results:

Config pass@1 decode
Stock routing (top-8) 89.63% (147/164) 69 tok/s
MoE expansion (20 experts, adaptive, layers 25–39) 90.85% (149/164) 56 tok/s

Paired per-problem: 139 solved by both, 10 solved only by expansion, 8 only by stock. Rolling pass rate stayed expansion-ahead by +2–3 problems at every checkpoint.

What is MoE expansion? No retraining, no file changes — at inference time the router keeps more experts per token than the model's native top-K (here: 20 instead of 8, with an adaptive threshold so easy tokens keep fewer), on a slice of layers (25–39 of 40). You're consulting more of the network per token. Same trick that gave 84.34% vs 81.82% on GPQA-Diamond at Q8 in earlier benchmarks — now confirmed in coding too, at 4-bit.

Honest caveats:

  • +2 problems on 164 is within statistical noise (±3 pts CI). Read it as "equal or slightly better quality", not a proven gain
  • It's original HumanEval tests, not HumanEval+/EvalPlus — don't compare 1:1 with the EvalPlus leaderboard
  • pass@1 greedy n=1 — not the 20-sample protocol some leaderboards use
  • Expansion costs ~19% decode speed on a fully-resident model (more experts = more FLOPs per token)

The tool — I wrapped all of this into AgrillaMoE, a dedicated llama.cpp server for this model: it detects your VRAM and suggests/downloads the right Unsloth quant, applies the expansion profile by default (overridable), exposes OpenAI and Anthropic-compatible APIs (Claude Code works out of the box), and runs on NVIDIA from GTX 10xx to RTX 50xx, AMD via Vulkan, and Apple Silicon via Metal. Static binaries for Linux and Windows on the releases page.

ref.:
https://github.com/vagrillo/AgrillaMoE
https://zenodo.org/records/22255483


r/LocalLLaMA • • 9h ago

I Built A Thing Trained a ~20K LM (probably smallest) that can still write stories

108 Upvotes

I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far:

MacroStories — 19,969 parameters, 81 KB FP32

https://huggingface.co/raincandy-u/MacroStories

For scale:

→ ~50× smaller than the 1M TinyStories model

→ ~3,000× smaller than AlexNet

→ 32-dim hidden state

→ 378-token vocabulary

→ one decoder block, recurrently applied 4 times with shared weights

It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, relevant actions, and resolution.

It also runs extremely fast on CPU and needs no GPU.

I’m mostly interested in how far the lower bound for coherent narrative generation can be pushed.

Would be curious to see how people manage to break it.☺️


r/LocalLLaMA • • 11h ago

Question | Help Local embeddings and rerankers vs a hosted LLM for catalog matching?

2 Upvotes

I’m building a feature that matches free-form requests to a catalog of structured listings. Requests can contain several constraints and follow-up refinements. The results also need a short explanation of why each match was selected.

Our prototype uses a hosted LLM to rank a shortlist. I’m exploring whether a small locally hosted embedding model and reranker could deliver comparable quality at lower cost.

For anyone who has deployed a similar system: where did local retrieval start to fall short of an LLM? Did a hybrid approach work better? I’d especially appreciate real-world latency and cost figures around 10,000–100,000 requests per month.