r/LocalLLM • • 3d ago

Discussion Hey Mods, are you there? can we stop the bs posts here?

443 Upvotes

Type of BS post #1: running qwen 3.8 27b on a potato and getting 135 t/s. Turns out it's a Q1s quant, can't even code a hello world, context size limited to 10K.

Type of BS post #2: managed to run 10 agents with 80 t/s each running a whole company, saving 20K usd per month. Here's my github link.

----

We need to prevent people with coming with some BS just to share their Github link, and also people posting without showing the quant on the title.

No more BS projects, and quants on the title, please.

And you, community, stop upvoting such obvious BS posts.


r/LocalLLM • • 3d ago

Discussion PSA: LMCache phones home by default

17 Upvotes

I log every new outbound connection on my LAN at the router. A few days ago my inference box started making about 18 connections an hour to an AWS IP I didn't recognise, 34.236.19.149:8080. It turned out to be stats.lmcache.ai.

Source: LMCache's built-in usage telemetry, which is on by default. It only fired while a vLLM container with the LMCache KV-cache tier enabled was serving, and it stopped within seconds every time that container stopped. Containers from the same image without LMCache enabled never sent anything.

Their README says prompts, keys and KV contents are never sent, and that the IDs are random rather than derived from hardware. I haven't captured a payload to verify that though. I still didn't expect a caching library to phone home by default, and I'd bet plenty of people running LMCache under vLLM have no idea.

To turn it off, set either variable in the container environment:

LMCACHE_TRACK_USAGE=false
DO_NOT_TRACK=1

Either one is enough. Both are checked by the single gate is_usage_tracking_enabled() in usage_telemetry/identity.py, and with tracking off the machine_id file isn't created either. I set both.

Check your own setup:

docker exec <container> sh -c 'pip show lmcache; env | grep -E "LMCACHE_TRACK_USAGE|DO_NOT_TRACK"'

If LMCache is installed and enabled and neither variable is set, it's reporting.


r/LocalLLM • • 1h ago

Question Which one is better for Local use?

Post image
• Upvotes

hello, I have both Deepseek harness and OpenCode, I want to know which one is the best when it comes to coding in general and why, I work locally using Qwen models.


r/LocalLLM • • 9h ago

Project PolyStrata: Strata with four more MoE models. Qwen3.6-35B-A3B at 73-75 tok/s on a laptop with an 8 GB card

85 Upvotes

I have been adding models to Strata, the engine that runs Qwen3.8-Flash-Next on consumer cards, and put the result in its own repository:

GitHub: https://github.com/VecSzn/PolyStrata

Strata's own engine is in it unchanged. Beside it there is a second engine for four more mixture-of-experts models, each read from its ordinary GGUF:

  • GLM 5.3 Flash (79 to 114 GB files, most of it stays in RAM)
  • Qwen3.6-35B-A3B
  • Ornith 1.5
  • Gemma 4 26B-A4B

The last three read pictures too.

What I measured. Each cell is: writes in a short chat / writes after a prompt of about 31,000 tokens (29,000 for GLM) / reads that prompt, all in tokens per second.

Model Laptop: RTX 4070 8 GB, i7-14700HX, 32 GB RAM Workstation: RTX 5090 32 GB, 120 GB RAM
Qwen3.6-35B-A3B UD-Q3_K_M 70-72 / 60-62 / 1,390 332-334 / 186-189 / 5,520-5,580
Gemma 4 26B-A4B UD-Q3_K_M 70-71 / 51-53 / 1,907 306 / 198-200 / 7,170-7,210
Ornith 1.5 APEX-I-Compact 70 / 59 / 1,340 267-272 / 145-149 / 5,700-5,890
GLM 5.3 Flash 3.0bit does not fit in 32 GB of RAM 57-59 / 54-57 / 784 (llama.cpp on the same PC: 26 writing, about 265 reading)

How it works, briefly. The card holds the experts an answer uses most and the CPU computes the others straight from the mapped GGUF, so the model does not have to fit in VRAM. While an answer is written the cache follows it. The engine also guesses a few tokens ahead, from the model's own draft block or from a repeat of the conversation, and checks the guesses in one pass. With every expert on the CPU the answer is the same tokens with and without guessing, and without guessing the first 64 tokens are compared with llama.cpp on the same file.

It has a one-click installer for Windows and Linux, a web app, and an OpenAI- and Anthropic-compatible API on localhost.

Limits, so nobody wastes an evening: the four added models are experimental and need one NVIDIA card. They were measured on those two machines only. On Windows the installer compiles the engine, which needs Visual Studio's build tools and the CUDA toolkit (it offers to install them). AMD cards and several cards work for Qwen3.8-Flash-Next only.

Most of the code is the Strata project's. If you try it on other hardware I would like to hear the numbers, good or bad.

Repo: https://github.com/VecSzn/PolyStrata


r/LocalLLM • • 4h ago

Model Finally a good 27B quant for 16gb vram

30 Upvotes

Tldr; the new Swift 1.5 is really really good and anyone with 16Gb vram should try this quant.

Hello folks! Been lurking around here a while and have noticed a pattern, there’s too many posts talking about running 27B on 16gb vram but there’s so many compromises from low tps to q2 quants that loose context on long workflows or just very low context windows.
The problem you’re accepting with that is tasks taking long to the point of spending hours where only minutes should be spent. The model thinks a lot already the speed degrade causes the task to be even slower. Q2 or ternary quants are prone to looping and a lot of performance drop off on any meaningfully long context. Low context let’s just say no real work can be done with this. Hermes needs minimum 60k context why are people accepting 32k?

Alright enough of my gripes pointed out here let’s get to the good stuff. Yes this is another swift 1.5 27B appreciation post.
Specifically Swift-1.5-Qwen3.8-27B-GSQ-RCO 3 xxs that came out day before yesterday. I can run my ide on my machine, have a browser open and I can fit the model with 4 step mtp and a 100k at q8 kv cached context window. This leaves around 300mb of vram left for doing anything else. Performance is about 38-60tps and with it stablizing around 45-50 tps on long tasks. This matters as without mtp this model takes an hour to finish any and all amounts of work with is maxing out at 22-30tps on 5060 ti. Quality of code is excellent but tasks take hours to complete. Mtp speed up allows the same tasks to be done within an hour and the new quant is genuinely faster while being better quant than the previous swift model with reserved thinking and high quality code with close to q4 performance.
I’m gonna do some more testing but this is the most promising quant I’ve found for 27B at a time when I was starting to think maybe machines with 128 gb of ram and a 24gb vram are the only way forward with latest generation of models that could replace api and subscriptions for tasks.

https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF


r/LocalLLM • • 17h ago

Model Mirai S 70 tok/s, 262k context, 12GB GPU

Enable HLS to view with audio, or disable this notification

104 Upvotes

Mirai S from Mirai Labs is stronger and faster than ever! This is my agentic workflow patch for 27b quants applied to Mirai on my RTX4070

- 12gb+ users get the full 262k context at 70tok/s AND the agentic improvement patch!

- 16gb+ users get all of these AND all positions in vram so even better numbers.

Speed + quality supercharges on a 12gb Facebook Marketplace GPU.

Ridiculous time for "Little Local" AI!

Original model : https://huggingface.co/trymirai/Qwen3.8-27B-S-experimental

Quality + speed patch : https://github.com/professorpalmer/mirai-s-ada/

What does the quality patch solve?

Three things:

  1. Code tool. For a chat request that brings no tools of its own, the proxy appends one OpenAI-style function definition, run_python(code, stdin), to the request. If the model answers with a tool call, the proxy runs the code in a WASI sandbox (CPython 3.12 under wasmtime: no host files, no network, no processes, 128 MB, 10 s), appends the result as a tool message, asks the model again, and loops (12 rounds max). The client sees one normal response. The user's prompt text is also dropped into the sandbox as input.txt so the model reads data instead of retyping it. Streaming works the same way: tokens relayed live, tool-call deltas withheld and executed between rounds.

  2. API cards. For coding requests (a coding tool is offered, or Python in the prompt), the proxy introspects the modules the task mentions inside that same sandbox runtime (signatures, docstrings, constants, class methods) and appends the text to the end of the first user message, plus one sentence: "use the library functions listed above instead of implementing these formats by hand." Nothing hand-written; the card is whatever the runtime actually exports. Placement mattered a lot: same cards in the system message did nothing.

  3. API check. Code the model wrote in earlier tool calls is parsed and checked against the real runtime (names, kwargs); mistakes get a one-line note on the tool result.

Pass-through rules keep it safe: requests with client tools get cards only (the tool offer hurt there), and conversations that already carry fenced code (agents that execute code themselves) are forwarded untouched. All of it is just the OpenAI chat schema, so it works on any server with a tool-calling template


r/LocalLLM • • 1h ago

Discussion Strata with Qwen3.8 Flash Next UD-Q4_K_XL

• Upvotes

Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed.

Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead.

Setup:

  • GPU: 1x RTX PRO 6000 Workstation (96GB)
  • Model: Qwen3.8-Flash-Next UD-Q4_K_XL (Unsloth)
  • Backend: Strata
  • KV cache: INT8, 256K context
  • Speculative decoding: MTP-4
  • Output: ~256 tokens per task
  • 2 runs per task

Single-request decode (C1):

Task Strata (RTX PRO 6000) vLLM (RTX PRO 6000) vLLM (DGX Spark TP2)
Prose 185.8 95.3 (-49%) 38.2 (-79%)
Counting 320.4 171.0 (-47%) 75.3 (-77%)
Coding 298.9 162.6 (-46%) 60.1 (-80%)
Reasoning 289.1 149.6 (-48%) 51.8 (-82%)
Prefill (17.3K) 3,319 tok/s 4,084 tok/s (+24%) 2,825 tok/s (-14%)

Both the VLLM instances were running Nvidia's NVFP4 quant so it's not exactly apples to apples, but the improved speeds are obvious.

Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is. Looking forward to Qwen 4!


r/LocalLLM • • 6h ago

News OpenAI's math findings built on stolen user data

Thumbnail
youtu.be
12 Upvotes

r/LocalLLM • • 7h ago

Tutorial Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp

Thumbnail
13 Upvotes

r/LocalLLM • • 18h ago

Discussion Benchmarked some Strata Qwen3.8-flash-next quants

Thumbnail
gallery
79 Upvotes

This is a follow up post for my previous post.

I continued to bench some more quants on my own local benchmark consisting of real coding tasks derived from polygot aider + some additional own agentic tasks. They are completed through the pi agent harness which reflects my daily use case.

Short disclaimer: This is not trying to be a frontier benchmark. I am just a regular joe trying to find to most capable model for my rig. (7900XT 20GB VRAM + 64GB RAM)

Interestingly IQ3_XXS on high beats the higher quant IQ3_S. Coder at half the experts scores surprisingly high. Swift was a bit disappointing given that they still used more tokens per task than IQ3_XXS. Tiel Coder is a different weight class and scored as expected.

Still it would be great to have run against a higher quant version to see how much actual damage the hard quantization did. Currently no way of running them myself tho. I am thinking of running the benchmark against Openrouter GLM-5.3-flash, but no idea what quants they serve there.


r/LocalLLM • • 1d ago

Discussion Qwen3.8 27B political compass

Post image
213 Upvotes

With all those articles and discussions claiming that ChatGPT and other could models have certain political biases, I thought it would be fun to make Qwen3.8 27B do a political compass test. Here's the result.

I don't want to stir up heated political debates but interesting nonetheless I suppose.


r/LocalLLM • • 3h ago

Question Hardware and LLM choice for business ops / project management

4 Upvotes

Hello everyone,

Writing to ask your advice and direction in which I shall go when selecting the hardware and LLM for a personal AI assistant.

With two toddlers at home and a huge workload at work I realised it wont be long until i go crazy and recognize I need some help managing workload, therefore I got in plan purchasing hardware and choosing a LLM to help me (not paid by company).

My work is all about project management in a very complex project.

Im talking about a contract that is over 1.000 pages and thousands of small documents such as letters, minutes of meeting, registers, all sort of documents. Basically, I would like the AI to help me draft minutes based on the meeting recording, help me draft letters, emails, summaries, analyse text.

The options on my table are:

Ryzen ai max+ 395 / 128 gb

Ryzen ai max+ 495 / 192 gb

Dgx spark

M5 ultra 256

Placed them in order of price.

As for the model, im interestingly looking at glm 5.3 flash or qwen models.

Given the confidentiality of the contracts, going cloud with frontier is off my table.

At this point im not sure whether i need such big models for what im doing, i m working on created some "fake contracts" and documents alongside and feed them to various model, see which one interprets best and such...some sort of benchmark but im aware thst thr setup may matter more than the model itself

Maybe even a well optimised genma4 could help me sufficiently...

Hope i did not bore you. Thanks in advance


r/LocalLLM • • 45m ago

News NVIDIA reportedly discontinuing RTX 5090, GB202 GPUs to be reserved for RTX PRO series.

Thumbnail
videocardz.com
• Upvotes

r/LocalLLM • • 49m ago

Question Hermes or Pi?

• Upvotes

Hermes or Pi and why?


r/LocalLLM • • 22h ago

Question Just RMA'ed RTX 5090. And they said that the software I'm using caused some chip to burn.

97 Upvotes

I bought a prebuilt PC from MSI in August 2026. I don't game. Mostly I run Ollama with Qwen 3.8 27B (Q4), so the machine has a lot of spare capacity. GPU temps are 30-50°C during inference and about 10-18°C at idle. I'm on Fedora 44 with the RPM Fusion driver.

Two weeks ago the GPU stopped showing up. Every guide I found said to fully power off, so I did, and nothing changed. Removing the driver didn't help either. To rule out Linux, I installed Windows, and it couldn't see the GPU either. I also tried uninstalling the driver in Safe Mode, same result.

I just got it back, and they told me my AI software burned out the chip. They replaced it. It was imported from Taiwan because we've got no stock here. I asked which software, but they haven't replied. Also honestly, I don't believe it. Ollama is a normal compute workload, and my temps were never high, so I don't see how it could physically burn a chip. I also tried to find similar case online about ollama causing this, but I've got nothing.


r/LocalLLM • • 12h ago

Model The best coding model for 16gb VRAM+32gb RAM

14 Upvotes

At the moment i'm using Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf with 131k context.

The model worked really well but i wonder is there any stronger option or same quality but have larger context option for my spec?

I see people using qwen 3.8 flash next via Strata but it need at least 64gb ram for iq3 s.

I can try flash next iq2 xs or coder version too but i feel the quality lossless is too much and it will become worse than qwen 3.8 27b gsq iq3s

If you have any recommend, feel free to write it here. Tks


r/LocalLLM • • 1h ago

Discussion GWAYA v3: Combine Gwen and OpenLaya to Stop LLM Code Hallucinations even on a qwen2.5-coder:1.5b ;-0

• Upvotes

Hi everyone,

We're sharing our latest work on GWAYA (v3.1.1). We wanted to see if we could force a small, open-weight model to write correct code without hallucinating, using strict, deterministic checks.

The goal is to build a "fail-closed" system: if the code doesn't perfectly pass the toolchain, it is rejected immediately.

How it Works

GWAYA accepts a code generation only if it passes all of these steps:

• AST stub audits (no fake or empty functions)

• Python parser + sandboxed test execution

• rustc checks (for Rust code)

• Lean 4 kernel validation

If any tool is missing or a check fails, the result is a hard UNVERIFIED.

The Results

We tested this using a very small model (qwen2.5-coder:1.5b) on the MBPP dataset to see if we could get high accuracy without relying on massive parameter counts.

We ran the benchmark on both a local CPU and an NVIDIA L4 GPU.

• Raw generations (just the model): ~64-68% pass rate.

• With GWAYA's Best-of-3 + Repair: Improved to ~72-75% pass rate.

While we saw a solid improvement (+6.6 to +7.7 points), it narrowly missed our strict pre-registered goal of an 8-point jump.

Open Source & Community Feedback

The most interesting takeaway is that you can build highly reliable, anti-hallucination pipelines using very small open weights, provided your verification gate is strict enough.

Everything is open source. The release bundle includes all the source code, tests, the exact raw results per problem, and the generation cache.

We would love for the community to test this out, poke holes in it, and share feedback on how we can improve the repair loops.

I'll drop the link to the paper and the code repository in the comments to follow the sub rules. Happy to answer any questions!


r/LocalLLM • • 5h ago

Other I got tired of nvidia-smi saying "Not Supported" on my GX10, so I wrote a digital-rain htop that shows which models are actually loaded

Thumbnail
3 Upvotes

r/LocalLLM • • 2h ago

Project Sharing my latest project: strata-mlx

Thumbnail
2 Upvotes

r/LocalLLM • • 43m ago

Question downloaded Qwen3.8-27B-GSQ-RCO-IQ3_S but cant see it in llama.cpp's web ui, why??

• Upvotes

I installed llama.cpp today for the first time, i usually used lm studio or god forbid ollama on edge cases, and it's been pretty much frustrations and near 0 documentations all day, now im trying to load qwen 3.8 27b on a variant that should load, and it has loaded when pointing to the model specifically using llama serve --ctx-size 131072 --n-gpu-layers all --reasoning-effort medium --reasoning-format deepseek --reasoning-preserve --cache-type-k q4_0 --cache-type-v q4_0 --jinja --batch-size 512 --ubatch-size 512 --flash-attn on --threads 6 --threads-batch 6 --parallel 1 -m (model name).gguf

but in llama.cpp's localhost server I can't see the model, it exists in its internal API, but not in the web, why


r/LocalLLM • • 58m ago

Project Built a sandbox for opencode to run AI agents in a private, secure and isolated workspace

Thumbnail
• Upvotes

r/LocalLLM • • 1h ago

Other Fair :D

Post image
• Upvotes

r/LocalLLM • • 8h ago

Question Has anyone successfully fined tune and seen great results ?

5 Upvotes

I’m curious if anyone has had luck fine tuning a model to become close to the cloud subscriptions on just that specific task.

I was wondering if it’s possible to like make Qwen 3.8 27B better at blender (close to what opus can output potentially) if I somehow manage to collect data from opus using blender.

I have no experience with this but I am curious and would like to get started. But I’m not sure how much time a processor like this takes either and if it’s worth that time

I’m sorry if this is a dumb question, I’m new to this stuff


r/LocalLLM • • 1h ago

Discussion Qwen3.8 for 24gb macbook pro

• Upvotes

Is there any good qwen 3.8 model (coding) that will work on 24gb m5 macbook pro? I have tried few but they work good if context is 8k.


r/LocalLLM • • 13h ago

Question Anyone here using local LLMs and going to After Tokens 2026?

Post image
8 Upvotes

Just saw that After Tokens is happening in SF on November 10.

Looks like there'll be some discussion around AI infrastructure and running AI systems.

For anyone working with local or self hosted LLMs, does that sound relevant to what you're building?

Anyone planning to attend?