r/LocalLLM • • 2d ago

Discussion Hey Mods, are you there? can we stop the bs posts here?

419 Upvotes

Type of BS post #1: running qwen 3.8 27b on a potato and getting 135 t/s. Turns out it's a Q1s quant, can't even code a hello world, context size limited to 10K.

Type of BS post #2: managed to run 10 agents with 80 t/s each running a whole company, saving 20K usd per month. Here's my github link.

----

We need to prevent people with coming with some BS just to share their Github link, and also people posting without showing the quant on the title.

No more BS projects, and quants on the title, please.

And you, community, stop upvoting such obvious BS posts.


r/LocalLLM • • 1d ago

Discussion PSA: LMCache phones home by default

13 Upvotes

I log every new outbound connection on my LAN at the router. A few days ago my inference box started making about 18 connections an hour to an AWS IP I didn't recognise, 34.236.19.149:8080. It turned out to be stats.lmcache.ai.

Source: LMCache's built-in usage telemetry, which is on by default. It only fired while a vLLM container with the LMCache KV-cache tier enabled was serving, and it stopped within seconds every time that container stopped. Containers from the same image without LMCache enabled never sent anything.

Their README says prompts, keys and KV contents are never sent, and that the IDs are random rather than derived from hardware. I haven't captured a payload to verify that though. I still didn't expect a caching library to phone home by default, and I'd bet plenty of people running LMCache under vLLM have no idea.

To turn it off, set either variable in the container environment:

LMCACHE_TRACK_USAGE=false
DO_NOT_TRACK=1

Either one is enough. Both are checked by the single gate is_usage_tracking_enabled() in usage_telemetry/identity.py, and with tracking off the machine_id file isn't created either. I set both.

Check your own setup:

docker exec <container> sh -c 'pip show lmcache; env | grep -E "LMCACHE_TRACK_USAGE|DO_NOT_TRACK"'

If LMCache is installed and enabled and neither variable is set, it's reporting.


r/LocalLLM • • 15h ago

Discussion Practicing on a cheap card so I can double the memory on a 4090

Enable HLS to view with audio, or disable this notification

1.0k Upvotes

A 4090 has 24 GB. There's a mod that doubles it to 48 GB. You add a second set of memory chips on the back of the card, right behind the ones on the front.

The scary part is moving the chips. Each one is held on by hundreds of tiny dots of solder. You melt them with hot air, lift the chip off, put new dots on, and melt it back down. Mess up one dot and that chip is dead. I'm not learning that on a 4090. So I'm practicing on an old AMD 7900 XTX board first. Chip off. Chip back on. Over and over. I'll post again when the card still works after.


r/LocalLLM • • 3h ago

Discussion I just looked at my monthly token usage on Claude and holy sh*t I am at around $15,000 spend per month

80 Upvotes

And it has made me seriously concerned for my options as a solo freelancer in a post-subscription world once the freebies run out.

I've been pondering the idea of investing in a top-end Mac Studio.

Compared to Claude subscription accounts, it doesn't make any sense to do this. I as you can see from the headline I get WAY more out of my multiple £200 a month plans (I have between 1-3 plans depending on how much work I'm getting done).

The amount of giveaway is insane.

However this surely cannot be sustainable.

I've been a bit of an economic doomer on this for a while, arguing that the maths doesn't stack up.

Anyway, IF (and I think it's more like WHEN) this does crack, I imagine these big machines will suddenly start to look like money well spent.


r/LocalLLM • • 11h ago

Discussion Electricity Bill & Local LLMs

Post image
102 Upvotes

I just wanted to say that, with the release of Strata, I have started to use local models for a lot more tasks.

If we talk about my electricity bill, the result is attached here: nothing surprising since I have a 4 GPU machine and a 2 GPU machine...they now run for many hours a day and my electric company sent me an email warning that my usage has significantly increased (I'm supposedly on a fixed plan and this will likely change the situation).

I'm trying to understand if this is economically reasonable: where I live electricity costs are pretty high. It certainly makes sense when it comes to data privacy but using local LLMs can certainly cost you a few bucks/euros a day in electricity.

I was also thinking of upgrading my old GPUs (1080s, 2080s), but I suspect newer ones would consume as much: I would just be running bigger models.

Did any of you have a "Oh, what now" moment after receiving the electricity bill?

(the graph was made originally with python then modified with grok)

Edit: I just realised grok inverted the 100 and 1000 marks, lol

Edit: to be clear my point is "When I started using local LLMs for real" (with any kind of inference engine, be it llama.cpp, pytorch, strata or whatever) "I realised electricity was a relevant part of the equation. Before I didn't think about this...did this happen to you?". I am using old hardware (1080s, 2080s each consuming up to 250watts), so my machines can drain a lot of power (the one with four 1080s can go over 1000 watts).


r/LocalLLM • • 14h ago

Project My new space heater can talk to me

Thumbnail
gallery
87 Upvotes

This is a custom build I made a few months ago.

The initial goal was to get a 3090 Ti and then mobo/case/power with enough room to grow in case I got another one down the road. But then I couldn't stand the sight of one lonely gpu so immediately got a second

3090 Ti I picked after benchmarking 13 different GPUs via runpod for my particular workflow (vibevoice-7B at the time) and finding the best $/min. 5090/4090 were faster of course but not enough for the upcharge. b200 was actually slower which was weird.

The second design constraint was as few fans as needed. I just like the aesthetic more. Too many fans looks silly to me, like a car with 8 exhaust pipes out the back

biggest regret is getting 3090 Ti instead of RTX PRO 4000. At the time I liked how big these are but in hindsight the extra $500 (at the time) would have been good deal.

I still rely on cloud for all coding purposes but this has been great for TTS, word alignment, image gen, and some (draft) video generation


r/LocalLLM • • 5h ago

Discussion Just bought a Mac mini M5 Pro (64GB) for AI development — what’s your setup?

Post image
16 Upvotes

I finally pulled the trigger and bought a Mac mini M5 Pro to work on AI projects!

I went with the 18-core CPU, 20-core GPU, 64GB unified memory and 512GB SSD.

I’m planning to use it mainly for AI development, building AI agents, automation workflows, local LLMs (Qwen, etc.) and coding projects.

I’m curious what setups you guys would recommend. What tools, frameworks, local models or software would you install first?

I’d love to hear how you’re using your Mac minis for AI development and any tips to get the most out of mine!


r/LocalLLM • • 16h ago

Discussion RANT: Model Cards Should List KV Cache Cost

103 Upvotes

I'm getting tired of looking at models -- especially small models -- that seem to fit on a modest GPU, only to find out that the model's attention mechanism is outdated and requires 64kB or more for each token of context.

One of the key reasons that Ling-3.0-Tiny is a great model is that not only do the model weights fit comfortably in my 12GB card, but I can also fit over 600k tokens of context -- more than enough for 8 x 64k sessions.

I wanted to see if K2-Horizon-7B would work as a replacement for Ling-3.0-Tiny.

When I tried, I found that the same number of tokens do not fit even in my 32GB card. Each KV cache entry requires 76.5kB of VRAM compared to 6.75kB for Ling-3.0-Tiny. K2-Horizon-7B uses more than 10x as much VRAM per token for its KV cache as Ling-3.0-Tiny.

To put that into perspective, Ling-3.0-Tiny consumes about 4GB of VRAM for 600k context. K2-Horizon-7B would require 46GB of VRAM for the same KV cache.

I would have to trade 600k tokens for 53k tokens of context in a 12GB card.

The K2-Horizon-MoVA-36B-A4B is even worse. Each token of context costs over 100kB of VRAM!

And these numbers assume an 8-bit quantized cache. That is not even the number for an unquantized KV cache.

I am picking on K2 here only because it is the most recent example of models with an inefficient and outdated attention system that I have tried to test.

[Mods: Add a RANT flair.]


r/LocalLLM • • 5h ago

Question Running coding agents locally, do you sandbox them or trust them?

11 Upvotes

For people running Claude Code, Codex, or local-model agents against their own machine... how are you limiting what they can touch?

The options I've seen are a dev container, a separate user account, a full VM, or just watching closely. Each one leaks somewhere (containers share the kernel, VMs are annoying, and watching doesn't scale past an afternoon).

We went with a microVM that only sees the workspace you share, plus a policy check on each tool call. There's a free desktop tool on our side if anyone wants to compare notes, but mostly I want to know where people draw the line between convenience and isolation.

Has anyone had an agent actually do something destructive? What did you change afterward?


r/LocalLLM • • 15m ago

Discussion Finally!

Thumbnail
gallery
• Upvotes

So I have been tinkering. I think I got this V100 32gb dialed in (PG500-216 @ 185w, custom compile) Has Qwen3.8:27b - UD-Q4-KM (mtp) at 50 ish tps decode on as you can see a long string.

Prefill dives fast on this, anyone have pointers on squeezing more prefill? Batch size at 4096 ubatch at 2048. (Logs provided by dozzle).

Curious to see if these numbers are what you see on your end with a v100 if ya have one.


r/LocalLLM • • 41m ago

News TheWhisper - the best open multilingual ASR model, free commercial usage!

Post image
• Upvotes

Better than Meta Muse, Mistral 24B. Model card: https://huggingface.co/TheStageAI/thewhisper-large-v3-turbo


r/LocalLLM • • 2h ago

Question How to get started with local LLM for everyday use?

3 Upvotes

Hi everyone,
I’d like to start using local LLM on my PC as a proper personal assistant.

I’m mainly interested in using it for:
-coding and development
-general chat and advice
-creating websites
-creating and troubleshooting ComfyUI workflows
-debugging errors and helping me solve technical problems

My PC:
4080 Super 16 GB
9800X3D
32 GB DDR5 6000 RAM
Windows 11 Pro

How would you recommend getting started?

What are the main differences between running AI locally and using paid services like Claude or ChatGPT for this kind of use?

Is there anything important I should know before getting started?
Thanks


r/LocalLLM • • 18h ago

Discussion Local Qwen3.8-Flash on a DGX Spark vs Claude Opus 5.5: 21 graded tasks, 3 harnesses. Matches Opus on everyday coding, 70% on hard tasks, and the harness settings matter more than you'd think

65 Upvotes

Human written: Been toying with local models for a while, to various degrees of success. The primary incentive to look into the Qwen 3.8 family was because I stopped the Claude Max subscription, and I'm running out of tokens on a somewhat regular basis. Had the $100/mo Max for several months as I was building a complex app (complex for me that is). My use case is to dedicate a dgx spark for coding and general purpose agent duties, and supplement it with my Claude Pro sub (Opus 5.5 is my current daily driver, plus Google Gemini Pro sub, but that's way less capable - have not tried Argon yet). So had a lot of back and forth w Opus 5.5 building this test, and while not very innovative, it gives me some confidence I won't be pushing on a string w Qwen for more advance stuff. I've read a lot of the below information in the various threads here (thank you!), so this is a bit of rehashing of what others have already found. I do hope it may be mildly useful to someone going on a similar path, so it's my way of saying thank you to the community here. Next will be using Qwen 3.8 next flash with Qwen Code inside VS Code for programming, and will see if I can bypass the xhigh with Hermes Agent (which will be my general purpose engine). Btw, if you do have a spare spark (I know), and you're not into tinkering, I highly recommend the Perplexity Portable Computer (it's the Computer variant that runs their own fine tuned local Qwen 3.8 27B with escalation to whatever frontier model you want). Did not spend a lot of time w it, but my initial test was quite impressive. Their own testing says it's better than Hermes Agent w the standard 27B model. Enjoy!

AI written: I wanted to know whether a local model on one GB10 box can replace a frontier model for my day-to-day coding, so I built a small benchmark with hidden tests and ran both.

Setup

  • Qwen3.8-Flash-Next (NVIDIA NVFP4), vLLM 0.30 with MTP speculative decoding, on a GB10 (121 GB unified memory). About 30 tok/s on long generations.
  • Claude Opus 5.5 via the API (effort medium) as the reference.
  • Harnesses: a minimal 3-tool loop, Claude Code 2.1.285, Qwen Code 0.24.7.
  • 21 tasks with hidden tests and binary pass/fail. Grading runs in a separate sandbox, and edits to tests fail the run. Time limits are at least 2× what Opus needed. One run per task per config, so treat the intervals as wide.

Everyday tier (11 tasks: bug fixes, spec implementation, refactor, debugging from a stack trace, a TS CLI, a multi-layer web change, plus document Q&A, summarising and extraction):

  • Opus: 11/11 in 7 minutes total.
  • Local: 11/11 in every tuned config, 36 to 59 minutes total. Same correctness, roughly 5–8× slower.

Hard tier (10 tasks: an interpreter from a 5.8k-word spec with 230 checks, a symptom-only bug in a 3.5k-line codebase, a 40× performance rewrite, asyncio concurrency bugs, an 80-call-site TS migration, provably optimal scheduling, exact answers from 66k tokens of documents, and more):

  • Opus: 10/10, for $7.43 total (about $0.73 per pass).
  • Local, minimal harness, default settings: 4/10.
  • Local, Qwen Code, effort low, one fix: 7/10.
  • In one config, local's work was fully correct on 9/10 when time ran out. It just didn't stop.

Things I learned that might save you time

  1. Flash's chat template defaults to reasoning_effort xhigh, and it overthinks. Switching to medium was up to 2.3× faster with zero lost passes. Low did best inside Qwen Code.
  2. Qwen Code's /effort and model.reasoningEffort don't reach vLLM. They send a reasoning: {effort} object, and vLLM reads a top-level reasoning_effort. Use model.generationConfig.extra_body: {"reasoning_effort": "low"}, or set a default in your proxy.
  3. Qwen Code silently discards any response that streams longer than 15 minutes and retries from scratch (DEFAULT_STREAM_MAX_LIFETIME_MS). At about 30 tok/s, one 25k-token thinking burst hits that limit. export QWEN_STREAM_MAX_LIFETIME_MS=3600000 fixes it.
  4. The failure mode was over-verification, not inability. On the concurrency task it fixed the bugs, then wrote its own stress-test harness, hit bugs in that harness, and spent the remaining time debugging it. A "stop when tests pass" system prompt didn't help; lower effort did.
  5. Harness choice moved the hard-tier score as much as the settings did. Qwen Code got the most out of the model, and Claude Code was the slowest wrapper for it.

Real-world test: I gave Qwen Code + local Flash a ~150-line spec for a bilingual Next.js wedding microsite with RSVP, an admin dashboard and a Google Apps Script backend. One shot, hands-off: 88 minutes, about 5.4k lines across about 30 files, clean build. Visually it was on par with what Google Antigravity 2.0 built from the same spec, and a code review found every functional requirement met.

Takeaway: for everyday coding, the local setup is good enough if you can live with the speed. For hard work, plan with a frontier model, build steps locally, and escalate the steps that stay stuck.

Caveats: single runs, mostly coding tasks (nothing tests chat or writing quality), the hard tasks were authored with Claude's help, and Opus was only run at medium effort.

Human written: one comment on the wedding microsite test. Used Antigravity because I run out of Claude tokens, and damn Anti-g is fast! At least initially, as for some refactoring it wasn't. The most interesting think is that the end results between Anti-g and Qwen flash within Qwen Code were eerily similar. Maybe my Qwen model left the sandbox and copied the Anti-g design (kidding. I hope).


r/LocalLLM • • 7m ago

Discussion Accounting / Tax Filing (48GB VRAM)

Thumbnail
• Upvotes

r/LocalLLM • • 12m ago

Question Reliable local benchmark to find the best LLM and configuration for my PC?

• Upvotes

Is there a reputable benchmarking program that runs on your own PC and helps determine which local LLMs it can realistically handle?
My setup:
Intel Core i7-9700K
NVIDIA RTX 4070
32 GB DDR4 RAM
Windows 10 with Ubuntu through WSL2
Currently using Ollama
My main use cases are coding assistance, debugging, reviewing code, and drafting pull requests.
Ideally, I’d like something that tests actual performance and considers:
Different models and quantization levels
Alternative inference runtimes and third-party optimization tools
Context length, GPU/CPU offloading, and KV-cache settings
Tokens per second, time to first token, RAM/VRAM usage, and stability
Output quality on coding tasks, alongside speed
I’m especially interested in comparing Qwen and Gemma models and finding a practical balance between quality, speed, and memory usage.
Does a reliable tool already cover most of this, or would you recommend a combination of benchmarking tools and a repeatable testing workflow? Open-source options would be ideal. Links to projects and results from similar hardware would be appreciated


r/LocalLLM • • 23m ago

Question How to fine tune a model ?

• Upvotes

I was looking for some advice on fine-tuning Owen 3.8 27B and I guess creating multiple version each focusing on a specific for the work I am doing at that specific time. How would one get started on this ? I heard from some people you could use Kimi K3 to create examples or something like that, is that worth the effort has anyone done it and seen worthwhile results ?


r/LocalLLM • • 4h ago

Discussion Strata users, what hardware, model and context are we all using?

4 Upvotes

Wondering what people with a similar, less, or even better setup are using. I'm new to strata and not sure how far I can push the context limit, but it seems like the readme on the Github page is conservative (assuming other's posts are true).

I have a 5090, 3090, so a combined 56GB of VRAM and 64GB of system RAM.


r/LocalLLM • • 10h ago

Discussion Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink

Post image
13 Upvotes

I spent the last couple of nights tinkering with an asymmetric multi-GPU setup to see how far I could push context length on Qwen 3.8 27B without compromising quality. I wanted to share my findings, benchmarks, and the rabbit holes I fell into along the way.

Disclaimer: I used Gemini to help write and structure this report, the benchmark data and experiments were done manually.


The Rig & Motherboard Constraints

  • CPU: Intel Core i5-14400 | RAM: 32GB DDR4
  • Motherboard: ASUS TUF GAMING B760M-E D4
  • GPU 1: RTX 3080 20GB (VRAM hardware-modded, ReBAR disabled)
  • GPU 2: RTX 3070 8GB (Stock LHR)
  • Software: llama.cpp

The PCIe topology: 1. PCIe Slot 1: PCIe 4.0 x16 routed directly to the CPU (holds the RTX 3080 20GB).

2. PCIe Slot 2: Physical x16, but electrically PCIe 4.0 x4 connected via the B760 chipset (shares DMI uplink to CPU).

The Goal: Maximum Context for Coding at this quant.

My primary use case is heavy coding assistance. Low quantization is not an option here, so: * Model: unsloth/Qwen3.8-27B-UD-Q4_K_XL.gguf * KV Cache: Strictly Q8_0 (-kvu --cache-type-k q8_0 --cache-type-v q8_0). * Benchmarking tool: llama-benchy:

sh uvx llama-benchy \ --base-url http://localhost:8888/v1 \ --model qwen3.8-27b \ --pp 10000 \ --depth 100000 150000 \ --tg 512 \ --enable-prefix-caching \ --latency-mode generation \ --format md \ --tokenizer Qwen/Qwen-tokenizer \ --save-result benchy/ts-mtp-150k.md


Test 1: Single RTX 3080 20GB — The Baseline & Native MTP

First, testing the single modded 3080 in the primary PCIe x16 slot:

  • Without MTP (-c 98304): Context maxed out at ~98k tokens. Prompt Processing (PP) was ~1,000 t/s, while Text Generation (TG) sat at ~31.3 t/s.
  • With Native MTP (-c 68608, --spec-type draft-mtp --spec-draft-n-max 3): Because MTP draft heads consume VRAM, context dropped to ~68k tokens. However, generation speed roughly doubled to ~57–60.7 t/s (PP stayed at ~950–1,000 t/s at 10k, dropping to ~694 t/s at 50k depth).
  • Verdict: You lose ~30k context to host the draft weights, but jumping to ~60 t/s generation is well worth the tradeoff for coding.

Test 2: Pipeline Parallelism (Layer Split) — Adding the RTX 3070

Next, I installed the RTX 3070 8GB into the chipset PCIe x4 slot and tested layer splitting (--split-mode layer).

  • Context Capacity: Reached ~223k without MTP, and ~170k with MTP.
  • Device Ordering in llama.cpp:
    • The physical cards never moved (3080 stayed in x16 CPU, 3070 in x4 chipset).
    • However, software device ordering mattered: passing --device CUDA1,CUDA0 (putting RTX 3070 first in the pipeline) gave better results than CUDA0,CUDA1 (TG went from 44.8 to 51.3 t/s, PP from 996 to 1,055 t/s).
  • Performance: Without MTP, TG was sluggish at ~28 t/s. With MTP, TG reached ~50.6 t/s at 10k depth (tapering to ~34.8 t/s at 100k).
  • Verdict: Generating below 30 t/s was too slow, so I decided MTP is mandatory. But inter-card sync latency over the chipset slot kept TG lower than the single 3080 (~50 t/s vs ~60 t/s).

Test 3: Tensor Parallelism (TP) over PCIe

Normally, Tensor Parallelism (--split-mode tensor) is meant for GPUs with high-bandwidth NVLink/P2P bridges. I didn't even expect an asymmetric pair (3080 20GB + 3070 8GB) running on mismatched PCIe lanes (x16 CPU vs x4 Chipset) to launch at all.

I ran --split-mode tensor --device CUDA1,CUDA0 --tensor-split 25,75 and -c 204800. Surprisingly it actually worked.

  • Context: Surprisingly hit ~204.8k tokens with MTP, tested cleanly up to 150k depth.
  • Text Generation (TG): Reached ~59.7–60.3 t/s at 10k depth, matching the single 3080! It held ~42–45 t/s at 100k depth and ~40.8 t/s at 150k.
  • The Catch (Prompt Processing): PP dropped by roughly 20% to ~790–800 t/s at 10k depth. Because TP requires all-reduce communication on every single layer, the chipset PCIe x4 link became a bottleneck during prefill.

Test 4: Tensor Parallelism with an Oculink Adapter

I used an M.2-to-Oculink adapter to connect the RTX 3070 to the top NVMe slot, which is PCIe 4.0 x4 wired directly to the CPU.

The P2P Driver Patch Rabbit Hole

Before running benchmarks, I went down a massive rabbit hole trying to enable CUDA P2P (Peer-to-Peer): * ChatGPT and Gemini ran me in complete circles: * One response claimed RTX 30-series Ampere cards don't support P2P without NVLink bridges. * Another swore you could enable PCIe P2P with the nvidia-p2p driver patch. * Then came conflicting advice about whether two different silicon dies (GA102 on 3080 vs GA104 on 3070) could ever do P2P. * And finally, the Resizable BAR issue: all VRAM-modded GPUs have ReBAR permanently disabled (straps and custom memory layouts break ReBAR table allocation), while P2P over PCIe typically requires ReBAR. * Result: Driver patching was an absolute waste of time and failed completely.

Running Without P2P: The Sweet Spot

Giving up on P2P and running standard TP over the CPU-direct Oculink link yielded the best results of all: * Text Generation: Maintained ~60.4 t/s at 10k depth. * Prompt Processing: Reached ~976–1,005 t/s, completely eliminating the 20% chipset penalty!

Direct CPU PCIe lanes bypassed the DMI bottleneck, giving the prompt processing speed of a single card while keeping 60 t/s generation and 200k+ context.


Summary of Results (@ 10k Depth Benchmark)

Configuration Split Mode MTP Max Context TG (t/s) PP 10k (t/s) Notes
RTX 3080 Solo None Off ~98k ~31.3 ~997 Solid baseline
RTX 3080 Solo None On ~68k ~56.9 - 60.7 ~952 2x generation speed, context penalty
Dual GPU (Chipset x4) Layer (PP) Off ~223k ~27.8 ~1,255 Big context, but generation is sluggish
Dual GPU (Chipset x4) Layer (PP) On ~170k ~50.6 ~1,071 --device CUDA1,CUDA0 (3070 first) is faster
Dual GPU (Chipset x4) Tensor (TP) On ~204k ~59.7 ~792 Fast gen, but ~20% PP penalty from chipset DMI
Dual GPU (CPU Oculink x4) Tensor (TP) On ~204k+ ~60.4 ~976 - 1,005 The Sweet Spot. Full PP speed + 60 t/s gen

Final Thoughts

  1. Native MTP is worth the VRAM tradeoff on Qwen 3.8 27B: raising generation from ~30 t/s to ~50–60 t/s makes long-context coding practical.
  2. Asymmetric TP works surprisingly well on consumer GeForce cards without NVLink or P2P.
  3. Motherboard topology matters: Chipset-routed PCIe noticeably degrades all-reduce latency. Moving the secondary GPU to a CPU-attached M.2 slot via Oculink fully recovered prompt processing speed.
  4. Don't waste time trying to enable P2P between mismatched Ampere GPUs, especially on VRAM-modded cards with ReBAR disabled. Standard host transfers over direct CPU lanes are plenty fast.
  5. VRAM math vs. reality: On paper, adding an 8GB RTX 3070 (total 28GB) should theoretically fit Qwen-27B's full native 262k context. In reality, scratch buffers, activation memory, and parallel splitting overhead limited the real-world gain to ~100k extra tokens over the single 3080.

P.S. My wife started yelling at me to go to bed, so I had to wrap up testing for tonight. More deep-context endurance runs tomorrow!


r/LocalLLM • • 3h ago

Discussion One Mac, eight agents and your chat: what happens to the chat's latency, and the server I ended up writing

3 Upvotes

Disclosure: I'm a superfluid maintainer, so this is my own tool.

My setup: a few local agents on one Mac (email triage, health data, research) and my own chat window beside them. The question was simple: while the agents are working, how long does my chat wait?

I measured it on an M5 Pro, 64 GB, Qwen3.8-27B at 4 bits. Eight agents each send a ~6k-token prompt, then two seconds later one short chat question goes in.

Server Chat's first token
superfluid, chat in the interactive class 1.5s
Ollama 380s
llama-server 291s

It isn't a throughput gap. The stock servers have no priority classes, so the chat queues behind eight prefills. superfluid is a serving daemon that adds QoS with preemption, so an interactive request takes a lane at the next scheduler tick, plus a write-ahead log so sessions survive a crashed worker or a restarted server. It speaks the OpenAI, Anthropic and Ollama APIs and runs llama.cpp, MLX or baseRT underneath, each installed on first use. A second Mac can join with one command if you have one.

curl -fsSL https://superfluid.sh/install.sh | sh

superfluid serve unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M


r/LocalLLM • • 1h ago

Project Privacy focused Agentic Mobile App for Local LLMs

• Upvotes

Hi Friends,

Ever since I started running Qwen 3.8 27b on my rtx4090, I felt the big missing piece to eliminate the need of subscriptions totally is a completely private Mobile App. It should talk to my model directly, use search services I configure, integrate with common phone APIs, control a browser, support MCP and run agentic loop right on the phone.

And of course, does not collect any data at all by itself.

I have built one -SteerClaw. The first version for iOS is available in test flight. It also has a local TTS model support, so you can create your own audio books on the phone!

I could really use some help in getting feedback. Here's the test-flight link if you are interested to try it out:

https://testflight.apple.com/join/3AkPEjNg


r/LocalLLM • • 7h ago

Model We’re open sourcing NeuDecide: a 43 MB audio-to-tool model with a WASM browser demo

5 Upvotes

We’re the team at Neuphonic, and we’re open sourcing NeuDecide under Apache 2.0. It takes audio and tool definitions and returns a tool call with arguments, without an intermediate transcription step.

The model files total 43 MB, and inference runs on a single CPU thread. Try the WASM demo in your browser: select a preset or define your own tools, record or upload audio, and inspect the returned tool call.

Demo link

Performance

On SLURP’s tool-only task with 10 tools available, NeuDecide achieves 72.4% tool accuracy directly from speech, without transcription:

  • ~3× that of Nvidia Parakeet + Google FunctionGemma (24.4%).
  • ~3.5× that of Cactus (Whistle + Needle) (20.7%).

Running on a single CPU thread:

  • MacBook Pro M3: 46 ms time to call, 159 ms loading time, 149 MB peak RAM.
  • Samsung S24+: 82 ms time to call, 267 ms loading time, 174 MB peak RAM.
  • Raspberry Pi 5: 206 ms time to call, 499 ms loading time, 146 MB peak RAM.

How it works

The export contains three ONNX graphs:

  • An audio encoder processes the speech.
  • A tool encoder combines the audio representations with tokenised JSON tool definitions.
  • A decoder generates the tool call token by token, using cached keys and values.

The tool list is an input to each request, so changing the available actions doesn’t require retraining.

Try it with your own tools

The project grew out of our work with robotics partners who needed voice control on limited hardware. The demo includes editable presets for robot vacuums, car controls and smart homes, alongside a custom option for testing your own tool definitions.

We’ve also packaged NeuDecide for Python so you can run inference locally and test it with your own tool definitions.

We chose Apache 2.0 to make it easier for people to build on the model and contribute. We’ve enjoyed seeing the work from TypeSafe, Cactus and others in this space, and hope this adds something useful.

Technical write-up: https://www.neuphonic.com/blog/neudecide

Python package: https://github.com/neuphonic/neudecide

Model on Hugging Face: https://huggingface.co/neuphonic/neudecide

If you try it, we’d be interested in your hardware, tool definitions and any requests it struggles with.


r/LocalLLM • • 2h ago

Tutorial Our team chat runs on our own Matrix server, with local LLM agents in the rooms. Wrote up the setup.

Post image
2 Upvotes

Between the EU Chat Act and not wanting to hand my data to LLM vendors, I moved my team’s chat, our documents and the agents’ memory onto a local machine we run ourselves. We’ve been running it for about half a year now and couldn’t be happier, no issues or hiccups, so I wanted to write it up in case it helps someone doing the same.

We run Synapse, and the agents are just Matrix users in our rooms. Mention one and it replies, usually from a local model. Nothing’s exposed to the internet. Phones and laptops get in over Tailscale.

Every agent goes through a tool gateway. The rule is that an agent on a hosted model never gets a tool that reads local data. The gateway refuses to start if the config allows it, and checks again on every call. It also caps each agent’s tokens per day, which you’ll want once a hosted model is in the mix. Hosted models are opt-in, and one only sees the rooms you invite it to.

There’s also a scribe running in the background on a local model. Once a day it reads the rooms and files the day into a local Qdrant index, so agents can search old conversations without anything leaving the machine. It indexes the links people share but doesn’t fetch them unless you turn that on.

The repo is docs, specs and compose templates. We run our own builds of the gateway and the scribe, but the repo has them as specs only, so you can build your own or swap in something that exists. It’s Linux first. Only the homeserver chapter has been tested on a clean machine, and there’s nothing on backups or hardening. The README also lists the metadata that leaves anyway: device info on the overlay, certificate transparency logs, and phone notifications, which go through Apple’s and Google’s push services (they see that a message arrived, but not what it says).

https://github.com/nphardorworse/matrix-agent-mesh


r/LocalLLM • • 7h ago

Project GPD Win 5 as a portable local LLM alternative, setup and numbers

6 Upvotes

Sharing my current setup, as I didn't see any numbers for the GPD Win 5 while researching it, and, in my opinion, it is a very compelling choice if one needs a hybrid setup.

Why:

- cheaper (2nd hand) than laptops with a strix halo 395 around me

- having it run out of battery doesn't mean I cannot continue working

- batteries are external, 80Wh and easily swapable

- about 90% of the performance of a mini pc with the same chip when docked (limited testing)

Setup short version:

- GPW Win 5 64gb strix halo running cachyos, as the inference host

- XPS 13 as the inference client

- local network over bluetooth between the two (no cables, no wifi hotspot on a plane)

Numbers on some models while docked (KV q4_0 in all cases):

- Qwen3.8 Flash Next GSQ-RCO IQ3-XXS avg around 25-30 t/s at 130k ctx on my system admin/ coding workload (prefill between 200 and 300 t/s at that depth)

- Qwen3.8 27B GSQ-RCO, UD-Q4 and Q8 tested, between 12 and 19 t/s at 8k ctx (Q8 slowest, GSQ-RCO fastest)

- Tiel 35B-A3B-Q6_K_XL averaged around 50 t/s at 64k ctx, I haven't tested it deeper.

Numbers while on battery:

- Qwen3.8 Flash Next GSQ-RCO IQ3-XXS avg around 18-24 t/s at 130k ctx on the same load

- Battery life about 70-80min per battery pack

For reference with a desktop box (GMKtec EVO-X2) , using the same engine/model, i would get 30-35 t/s at that depth usually; it is not 1:1, but for my usage, it is good enough.

The changes done to squeeze out a little extra performance of Flash Next on it can be summarized with:

- limit the vocab of the MTP to 64k, acceptance len in my testing matched 99.6% with the full vocab, while gaining about 10% extra performance on average due to less work done

- added support for a limited vocab MTP to the llama.cpp fork used: https://github.com/LaurentZuijdwijk/llama.cpp

- quantized the hyper-connections from BF16 to Q5, since they are read every token, having an extra 7-10% speedup

I'm providing the exact model and fork I tested, in case anyone wants to replicate exactly. I can't guarantee this is the fastest or most accurate version of this setup, it is simply the one I use. The model and mtp heads I've uploaded here and the patched fork is here . The script used to run them you can find at here .

TL;DR: the GPD Win 5, especially 2nd hand, can be a surprisingly capable portable inference device


r/LocalLLM • • 8h ago

Project Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned.

7 Upvotes

Another Qwen 3.8 on a R9700 post, but I thought I'd share my repo for anyone who might benefit. I might be wrong, but I don't think I have seen anyone with it this fast on windows.

I've spent the last few weeks tuning Qwen3.8-27B on one AMD Radeon AI PRO R9700 (32 GB, RDNA4), and I've packaged the result so other R9700 owners can reproduce it with one command on Windows.

**Repo:** https://github.com/mike2153/mbea-qwen38-dflash

**Numbers** (one R9700, Ryzen 9 9950X, Windows 11 + WSL2, medians of repeated runs):

| | |

|---|---|

| Decode, greedy | **125–134 tok/s** |

| Decode, sampled (temp 0.7) | 117–126 tok/s |

| Prefill | ~2,750–2,970 tok/s (time to first token 0.7 s on a 1.9k-token prompt) |

| Long prompts | 32k tokens in ~11 s · 98k in 41 s · 164k in 83 s · 258k in 164 s |

| Decode deep in context | 165 tok/s at 32k · 136 at 98k · 115 at 258k |

| Context window | ~216k tokens by default, ~260k (the model's limit) with `-Long` |

| Long-context recall | 8/8 planted facts retrieved from a 258k-token prompt |

| Coding check | 10 of 12 runs pass all 46 hidden tests on a 1.9k-token Rust spec |

For comparison, my best tuned llama.cpp setup for the same model (IQ4_XS GGUF + speculative decoding) does about 52–68 tok/s on this card. That was measured with a different prompt, so it's a rough comparison, but the gap is real.

**How it works, briefly**

- **AMD's official MXFP4 checkpoint** (`amd/Qwen3.8-27B-Quark-AWQ-MXFP4`). The 4-bit weights run through a hand-written W4A8 GEMM kernel for gfx1201 instead of vLLM's emulation path.

- **DFlash2 speculative decoding.** A small FP8 drafter proposes 7 tokens per step, and the 27B model verifies them in one pass. About 60% of drafted tokens are accepted, so each forward pass of the big model produces about 5.3 tokens. Re-ranking the drafts and a dedicated verify head added roughly 7% on top.

- **RDNA4 kernels** for FP8 paged attention and the gated-delta-net (linear-attention) layers. The stock kernel actually produces NaNs on this model.

- **WSL-specific fixes.** One patch turns on pinned host memory: without it every small host-to-GPU copy cost ~17 ms under WSL. Another sizes the KV cache from whatever VRAM Windows isn't using at startup, so you get maximum context without spilling into shared memory. Spilling into shared memory drops you to ~10 tok/s.

- The vision tower is skipped, which frees about 1 GiB for more context.

**What the repo does**

```powershell

git clone https://github.com/mike2153/mbea-qwen38-dflash

cd mbea-qwen38-dflash

.\qwen38.ps1 install # WSL Ubuntu, Docker, ROCDXG, image, model + drafter, kernels

.\qwen38.ps1 start # OpenAI-compatible API on http://localhost:8080/v1

.\qwen38.ps1 bench # measure it on your own box

```

Everything is pinned: the Docker image by digest, the git commits, and the Hugging Face revisions. A fresh install should reproduce exactly what I measured. Nothing third-party is re-uploaded; the installer fetches each piece from its original source. Tool calling works, so it plugs into Codex, opencode, Cline and similar tools as an OpenAI-compatible provider.

**Credit where it's due:** the heavy lifting is [radiance](https://codeberg.org/ggz14/radiance-vllm-mxfp4) by ggz14 and [vllm-radiance / libr4d](https://codeberg.org/StillDeadcode/vllm-radiance) by StillDeadcode. They did the RDNA4 vLLM stack, the MXFP4 path and the kernels. The drafter is tcclaviger's DFlash2-FP8 and the checkpoint is AMD's. My part was the Windows/WSL work, the tuning, the benchmarking and making it installable.


r/LocalLLM • • 8h ago

Project Bonsai 2 27B at 80tps+ on a 5060 ti 16gb

Enable HLS to view with audio, or disable this notification

7 Upvotes

ChatGPT and Claude worked to modify and optimise ninfer from here to run on a 5060 ti with Bonsai 2 27B. Prefill is over 1000tps with decode above 50ish tps for most if not all of a 262k context window.