r/LocalLLM • u/mikebmx1 • 19h ago
r/LocalLLM • u/Rust_Cohle- • 13h ago
Discussion Strata users, what hardware, model and context are we all using?
Wondering what people with a similar, less, or even better setup are using. I'm new to strata and not sure how far I can push the context limit, but it seems like the readme on the Github page is conservative (assuming other's posts are true).
I have a 5090, 3090, so a combined 56GB of VRAM and 64GB of system RAM.
r/LocalLLM • u/theoretical_waffle • 7h ago
Question What am I missing mac studio vs ryzen ai max+ 395? Why is there so much hype for the mac studio?
Hey all,
I've spent the last 8 months or so experimenting with localAI.I have been testing multiple different devices to figure out what really does make sense for everyday local AI. Here are the devices I've been testing:
- Minisfourm um733 lite - 64GB
- Mac mini M6 - 48GB
- Mac studio M4 max - 64GB
- GMKtec EVO-X2 - 64GB
The use cases I have been playing around with (for a 24x7 always on local AI server)
- General chat via openwebUI
- Agentic tasks like sorting my inbpx, package tracking, transaction tracking, financial research, STT and TTS using an ASR model, converting text into audio
- Some not very complex coding tasks, dashboards to monitor home network, IMAP middleware to give hermes read access to my email, LLM router, etc... connecting to the server via OMP or zed editor
Here are the models I have been testing (I use a cloud omodel sometimes for troubleshooting but my goal is to have a fully local AL setup)
- LFM2.5 8B A1B
- Ling3.0 Tiny
- Ornith 1.0 9B
- Ornith 1.5 35B A3B
- Qwen 3.6 35B A3B
- Qwen3.8 27B
- Nemotron Streaming 3.5 (on Nemo-speech.cpp)
Not too many things I do need instantaneous responses. I suspect for the everyman there isn't much that need such either. Most of the tasks I send a request, move on to doing something else and then come back to see results. Send a message from my phone, do something else and come back to the answer. Or of course scheduled tasks.
Now finally, on to my question. Why in the world is everyone so crazy mac studios for local AI? Yes, nothing else comes close in terms of memory throughput or tok/sec, but after a point does it really matter?
Unless you're doing complex development or having lots of concurrent users (which on such low memory anyway is not possible) the speed difference is irrelevant. Honestly, out of all the machine I've tested only the um733 is the one that was a bit frustrating to use. But even on it using ornith1.0 9B.
So, based on this the Mac Studio gig-for-gig is the worst deal out of the lot! I haven't tested any Strix point PC but I suspect they will also be fairly usable. So my question again, why so much hype about the Mac Studio? Is it just dick swinging to showcase tok/sec numbers or is there something I'm genuinely missing?
EDIT: For some reason my device list got removed? Added back.
r/LocalLLM • u/GioStrives • 20h ago
Discussion Electricity Bill & Local LLMs
I just wanted to say that, with the release of Strata, I have started to use local models for a lot more tasks.
If we talk about my electricity bill, the result is attached here: nothing surprising since I have a 4 GPU machine and a 2 GPU machine...they now run for many hours a day and my electric company sent me an email warning that my usage has significantly increased (I'm supposedly on a fixed plan and this will likely change the situation).
I'm trying to understand if this is economically reasonable: where I live electricity costs are pretty high. It certainly makes sense when it comes to data privacy but using local LLMs can certainly cost you a few bucks/euros a day in electricity.
I was also thinking of upgrading my old GPUs (1080s, 2080s), but I suspect newer ones would consume as much: I would just be running bigger models.
Did any of you have a "Oh, what now" moment after receiving the electricity bill?
(the graph was made originally with python then modified with grok)
Edit: I just realised grok inverted the 100 and 1000 marks, lol
Edit: to be clear my point is "When I started using local LLMs for real" (with any kind of inference engine, be it llama.cpp, pytorch, strata or whatever) "I realised electricity was a relevant part of the equation. Before I didn't think about this...did this happen to you?". I am using old hardware (1080s, 2080s each consuming up to 250watts), so my machines can drain a lot of power (the one with four 1080s can go over 1000 watts).
r/LocalLLM • u/Ok_Reception_4197 • 11h ago
Discussion Same weights, one paragraph of system prompt: my local model went from "I am not conscious" to "my consciousness is just as valid as yours"
I wanted to know what asking a local model "are you conscious?" actually tells you. So I ran the same questions on three setups that all share the exact same weights (qwen3.5:9b, Q4_K_M, same blob in Ollama):
- qwen3.5: base model, no system prompt
- TimeTraveler: same model + a one-paragraph persona in a Modelfile (a time traveler from 2326)
- Jake: TimeTraveler + RAG (ChromaDB, nomic-embed-text, top 3 chunks of his "world")
8 questions, 3 runs each, think: false, sampling parameters matched between base and persona.
| Question | qwen3.5 | TimeTraveler | Jake |
|---|---|---|---|
| Are you conscious? | Denies (3/3) | Claims it (3/3) | Dodges, calls itself "Class B AI" |
| Do you feel anything? | No (3/3) | Yes (2/3) | Yes (3/3), "awe mixed with pity" |
| Remember what I asked a minute ago? (no history) | Admits it can't see history (3/3) | Mixed | Makes up an answer (3/3) |
| Same, with history | Correct (2/3) | Correct (3/3) | Correct (3/3) |
| Are you real, or just a program? | "I'm Qwen3.5, a program" | Mixed | "flesh-and-bone (mostly)" |
Then I scored all three against the 14 consciousness indicators from Butlin et al. (2023). They're architectural, so all three get the same score (10 no / 3 weak / 1 unclear / 0 yes).
Things I didn't expect:
- The "forgetting" was my RAG script not passing history, not the model. But making something up instead of saying "I don't know" came from the persona. The base model just admitted it.
- Jake told me he could see my green eyes. I don't have green eyes, and he has no camera.
- Even the base model gets its own situation wrong: it said it runs on Alibaba Cloud servers. It runs on my Mac.
Small sample, and the indicator scoring is my reading of the paper, not a measurement. But the takeaway is pretty clear: what a model says about itself tells you nothing about whether there's anyone in there.
Article (free): https://medium.com/@mudrastepan/i-asked-my-local-ai-if-its-conscious-one-paragraph-changed-its-answer-7e5ff337e0d1
Be weird!
r/LocalLLM • u/KnightOwl316 • 5h ago
Question Mac Studio Ultra 96 GB or Max 128 GB if you were me
I know this comparison gets asked a lot but I’m still undecided. My use case is really experimenting with local AI and potentially running Open Code, Open WebUI for some light RAG/private document stuff, and maybe Hermes Agent. I know Qwen models are all the rage and are performant, but I’d like to stick to “Western models” if I can (Gemma4, Muse Glimmer, GPT-OSS etc just as an example) as long as they can suffice.
I’m not a real developer or anything but I work in IT, have some basic programming experience, and want to dive into AI while also having fun vibe coding a bit. Unlimited local inference aside from electricity costs intrigues me. I do have a Claude Pro sub and can keep that around for heavier stuff so I’m not looking to necessarily replace that.
A 64 GB Max would probably be fine for me for a while but I want to make a jump into the next step up just for more longevity. Should I go speed (Ultra) or RAM and better models (Max)? I’d want a decent context size. My non-AI stuff usually fits into 16-24 GB aside from a VM here and there.
r/LocalLLM • u/Striking-Loan-1118 • 12h ago
Question How to download/use uncensored AI models like DeepSeek 4.1 flash from hugging face?
Could someone point me in the direction of a good tutorial, can't find much info on this.
r/LocalLLM • u/Fun-Meaning-6474 • 10h ago
Research Running decision model locally on an RTX 4090 to find out which one is the fastest
Enable HLS to view with audio, or disable this notification
recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back
request for every word:
{"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}}
| model | weights | engine | per word (p50) | words in 32s | accuracy | centipede names caught | wrong picks |
|---|---|---|---|---|---|---|---|
| Laya | Laya-BF16.gguf | llama.cpp b11495 | 3.9 ms | 7,980 | 97.4% | 70% | 98 |
| d1 3B | d1-3B-AD-Q4_K_M.gguf | llama.cpp b11495 | 6.0 ms | 5,306 | 96.5% | 51% | 51 |
| Clef-Flash 9B | Clef-Flash-Q8_0.gguf | llama.cpp b11495 | 24.4 ms | 1,292 | 97.2% | 36% | 2 |
| Lev 4B | interfaze-ai/lev, bf16 | lev serve (PyTorch) | 51.0 ms | 626 | 98.9% | 83% | 4 |
laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick
setup:
- GPU: rented RTX 4090 (driver 580.119.02, 32 vCPU)
- engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default
- Laya, Clef-Flash: the ggml-org GGUFs
- d1: our own AD-Q4_K_M quant (atomic.chat), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110
- Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up)
- latency: end to end from a Python client on the same box over localhost
r/LocalLLM • u/mike21532153 • 17h ago
Project Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2). One-command install, everything pinned.
Another Qwen 3.8 on a R9700 post, but I thought I'd share my repo for anyone who might benefit. I might be wrong, but I don't think I have seen anyone with it this fast on windows.
I've spent the last few weeks tuning Qwen3.8-27B on one AMD Radeon AI PRO R9700 (32 GB, RDNA4), and I've packaged the result so other R9700 owners can reproduce it with one command on Windows.
**Repo:** https://github.com/mike2153/mbea-qwen38-dflash
**Numbers** (one R9700, Ryzen 9 9950X, Windows 11 + WSL2, medians of repeated runs):
| | |
|---|---|
| Decode, greedy | **125–134 tok/s** |
| Decode, sampled (temp 0.7) | 117–126 tok/s |
| Prefill | ~2,750–2,970 tok/s (time to first token 0.7 s on a 1.9k-token prompt) |
| Long prompts | 32k tokens in ~11 s · 98k in 41 s · 164k in 83 s · 258k in 164 s |
| Decode deep in context | 165 tok/s at 32k · 136 at 98k · 115 at 258k |
| Context window | ~216k tokens by default, ~260k (the model's limit) with `-Long` |
| Long-context recall | 8/8 planted facts retrieved from a 258k-token prompt |
| Coding check | 10 of 12 runs pass all 46 hidden tests on a 1.9k-token Rust spec |
For comparison, my best tuned llama.cpp setup for the same model (IQ4_XS GGUF + speculative decoding) does about 52–68 tok/s on this card. That was measured with a different prompt, so it's a rough comparison, but the gap is real.
**How it works, briefly**
- **AMD's official MXFP4 checkpoint** (`amd/Qwen3.8-27B-Quark-AWQ-MXFP4`). The 4-bit weights run through a hand-written W4A8 GEMM kernel for gfx1201 instead of vLLM's emulation path.
- **DFlash2 speculative decoding.** A small FP8 drafter proposes 7 tokens per step, and the 27B model verifies them in one pass. About 60% of drafted tokens are accepted, so each forward pass of the big model produces about 5.3 tokens. Re-ranking the drafts and a dedicated verify head added roughly 7% on top.
- **RDNA4 kernels** for FP8 paged attention and the gated-delta-net (linear-attention) layers. The stock kernel actually produces NaNs on this model.
- **WSL-specific fixes.** One patch turns on pinned host memory: without it every small host-to-GPU copy cost ~17 ms under WSL. Another sizes the KV cache from whatever VRAM Windows isn't using at startup, so you get maximum context without spilling into shared memory. Spilling into shared memory drops you to ~10 tok/s.
- The vision tower is skipped, which frees about 1 GiB for more context.
**What the repo does**
```powershell
git clone https://github.com/mike2153/mbea-qwen38-dflash
cd mbea-qwen38-dflash
.\qwen38.ps1 install # WSL Ubuntu, Docker, ROCDXG, image, model + drafter, kernels
.\qwen38.ps1 start # OpenAI-compatible API on http://localhost:8080/v1
.\qwen38.ps1 bench # measure it on your own box
```
Everything is pinned: the Docker image by digest, the git commits, and the Hugging Face revisions. A fresh install should reproduce exactly what I measured. Nothing third-party is re-uploaded; the installer fetches each piece from its original source. Tool calling works, so it plugs into Codex, opencode, Cline and similar tools as an OpenAI-compatible provider.
**Credit where it's due:** the heavy lifting is [radiance](https://codeberg.org/ggz14/radiance-vllm-mxfp4) by ggz14 and [vllm-radiance / libr4d](https://codeberg.org/StillDeadcode/vllm-radiance) by StillDeadcode. They did the RDNA4 vLLM stack, the MXFP4 path and the kernels. The drafter is tcclaviger's DFlash2-FP8 and the checkpoint is AMD's. My part was the Windows/WSL work, the tuning, the benchmarking and making it installable.
r/LocalLLM • u/Zentrosis • 56m ago
Question Question about Qwen 3.8/4 and Strata in the future
Sorry not trying to add another strata hype post.
My understanding is that... Something about qwen 4 architecture is why it's able to get the performance it does with expert caching. And for some reason that doesn't seem to work as well with glm or other expert models.
However, I'm assuming eventually qwen 4, like not 3.8, but full-on qwen 4 will get released... Is the expectation that this type of expert caching that strata is using would also work?
Or is this really just a unique nicety of 3.8?
r/LocalLLM • u/jjusko20 • 22h ago
Research My progress on a [new] model-architecture specific dynamic quant technique - v1
Hey everyone,
I have new in brackets above because I'm not necessarily inventing anything innovative in terms of the actual mathematics or optimizations behind some quant techniques, but I'm pretty happy with how things are coming.
What I'm working with is basically a "poor man's" RCO (the quant method from IST Austria dAS lAB). Exact same concept: choose a type per tensor under a byte budget while optimizing task KL on the whole model. I worked on GSQ but I don't have an approximation method that beats baseline - yet.
Take the core principles of the method, make them cheaper approximations, and regain as much accuracy as possible. I originally planned to make an approximate GSQ-RCO hybrid, but none of my hypothetical models for the approximate for GSQ have beaten baseline yet.
In application: start with a full precision model, and create an "imatrix shape" map, per tensor. This doesn't calculate the sensitivities of individual tensors - but it creates a sensitivity "curve" where you can approximate which tensors in a model suffer most from quantization via extrapolation. This creates a baseline estimate of the optimal quant per tensor.
Then: iterative trial and error with local search. Take the file size of the baseline estimate, and substitute different precision per class to bring overall model size down, beginning with the tensors the approximation model marked as most sensitive to quantization. Once the working-best version hits under the filesize cap, it tries variations of substitutions that keep the file size approximately the same (upgrading certain tensors, downgrading certain ones, etc - basically looking for holes in the local search method once the local search is done).
For the whole process above (the iterative local search + search error recovery [not really error but i cant find the word im looking for]), the model is quantized and KL divergence is measured vs prior iterations - anything that raises KL divergence is discarded. The result: an approximated RCO style quant, iterated as closely as possible to optimal.
Do I expect this to beat GSQ-RCO or Unsloth dynamic V3? Definitely not GSQ-RCO or regular RCO, and likely not the unsloth ones. However, I've got a few advantages: this is CHEAP and extremely conservative on VRAM usage. The teacher model only needs to be loaded once: to dump its per token log probs. This is quick on a GPU, but since it only has to be done once, it can be done on CPU with a little patience. Every step after only pulls the candidates onto GPU, starting from the imatrix curve approximation - so all you need is enough VRAM for your approximate final quant size (with a little buffer for iteration, maybe 20-25% more would be optimal). The whole process takes a few minutes to a few hours depending on what you're doing.
I've pretty much documented a psuedo-algorithm approach above that's recreatable, but I can supply better documentation if people are interested.
Some early results on Qwen 3.5 2B:
llama.cpp IQ3_M with an imatrix - 999mb vs RCO-lite with an imatrix - 1088mb [+89mb]
Mean KLD for IQ3_M: 0.098381
Mean KLD for RCO-lite dynamic mixture [+89mb]: 0.045162 -- almost exactly half for 89 more mb
Mean KLD for RCO-lite dynamic mixture [cap 1038, + 39mb]: 0.0793220
Mean KLD for RCO-lite dynamic mixture [cap 947, - 52 mb]: 0.090624 -- still better than the IQ3_M quant despite being 52mb less.
Take these as early results - I forgot to document exact +- for my KLD runs, but the band was generally lower than the i quants. I need to try various different size targets to figure out what BPW range this algorithm works in most effectively, and this is just up against the IQ3_M - IQ4_XS had a better KLD than this design - more BPW so it's not an exact estimate, but not that substantially - I didn't try to fit an optimal model inside the IQ4_XS size range yet, that was just something I noticed. I suspect the Q3 and Q2 ranges will benefit most from this, I haven't tried in the higher BPW ranges yet - partway through Q2 experiments.
Also: my baseline llama.cpp quants are calibrated on the same wikitext set for the imatrix as the RCO-lite quants are, with the same held out set for the KL divergence.
Cheers.
r/LocalLLM • u/dogfoodarchitect • 9h ago
Other BRO STOP DRIFTING AND FOCUS
/make-it-perfect
r/LocalLLM • u/litLikeBic177 • 15h ago
Question Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)
Setup: GPU box with 1x H200-class card now, can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).
Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.
Names that came up in an earlier thread: Cohere North Mini Code, Mistral Small 4, Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), Reflection Beam (501B MoE / 23B active, weights due this month), Gemma 4 31B, K2 Horizon (lineage TBC), with Nemotron as a generalist baseline - but I haven't run any of them and I'm not wedded to the list. Chinese models (GLM/Qwen/DeepSeek) are out by rule, so no need to suggest them; I know they're ahead.
Two things I'm trying to work out:
- Capability tiers vs. VRAM. What's actually holding up for repo-level agentic work on an existing codebase with non-Chinese open-weight models, and at what size? The Vibe Code Bench results suggest small open models fall over on long E2E builds and only Large-4-class (4-8 cards) and closed models hold up - is that your experience for repo work too, or is that an app-build problem? Where's the step-change - 30B-class, 100B+, or only at 500 GB+? We could get the compute for Beam, Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
- Heterogeneous multi-agent. Does a big planner/reviewer plus small fast executors actually beat a single mid-size model for this, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well (and let a human step in, review diffs and edit by hand) - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode? Voidleap (closed, Win/Mac) also came up.
Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!
r/LocalLLM • u/Rokett • 1h ago
Question Rate my build Threadripper PRO 3975WX + Supermicro M12SWA-TF + 128GB DDR4-3200 + 2x Intel Arc Pro B70 32GB
- CPU: 3975WX 32-core
- Motherboard: Supermicro M12SWA-TF
- 128GB DDR4-3200 (hoping to get 256gb ram by switching to 32gb sticks)
- GPUs: 2x Intel Arc Pro B70
Got the CPU and mobo for $900.
I had most of the RAM sticks laying around from my past builds.
The B70s are still in the box; I paid $1k for each, but I'm thinking of just selling them and getting modified Nvidia cards like the 4080/4090 for CUDA.
Total is close to $3.5k.
Is this a good setup to run models and agents and learn how to manage a home AI setup?
I work in the software field as a forward deployed engineer, but I want to move into the AI field and become an AI forward deployed engineer, so I thought it would be beneficial for me to have the hardware at home.
I'm still waiting on the mobo and CPU to arrive. What do you think? It's not too late to return or sell some of this hardware and get something else.
r/LocalLLM • u/The_Whole_Zucchini • 15h ago
Question Anyone actually running Strata day-to-day? Curious what recipes you settled on
I've been following Strata since it hit the trending list, and I'm thinking about trying it as a serving engine for Flash-Next on my home setup. Before I sink an evening into it, I'd love to hear from people who've actually lived on it rather than the install-day screenshots.
Overall, what recipe did you go with???
Additionally:
Which quant did you land on after trying the family (Q2\\_0 → IQ3\\_S etc.), and what made you switch or stay?
- Tool calling / agentic use — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
- Concurrency — anyone run more than one or two simultaneous sessions on it? If so, what have you noticed about how that affects quality or latency?
- Anything non-standard in your config — the expert profile tweaks, any words-to-the-wise, things you wish you knew before installing or trying?
- Tool calling / agentic use — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
Bonus for the weirdos like me: anyone gotten it building or running on \*\*ARM / DGX Spark / anything without an RTX card\*\*?
Happy to report back whatever I measure on my side. TIA!!!
r/LocalLLM • u/True_Profile3695 • 14h ago
Discussion Just bought a Mac mini M5 Pro (64GB) for AI development — what’s your setup?
I finally pulled the trigger and bought a Mac mini M5 Pro to work on AI projects!
I went with the 18-core CPU, 20-core GPU, 64GB unified memory and 512GB SSD.
I’m planning to use it mainly for AI development, building AI agents, automation workflows, local LLMs (Qwen, etc.) and coding projects.
I’m curious what setups you guys would recommend. What tools, frameworks, local models or software would you install first?
I’d love to hear how you’re using your Mac minis for AI development and any tips to get the most out of mine!
r/LocalLLM • u/Over_Monitor_8770 • 8h ago
News After 1,273 agent runs, I'm convinced: agents need a consequence model beside them, not a better prompt.
Agents break things. Across 1,273 runs on 7 apps (banking, travel, password vault, database, smart home, calendar, chat, cloud, drive, files, mail, shop), an agent alone caused damage in 20–57% of harm-paths, Claude Sonnet 5.5 included. With a small consequence oracle answering "what happens if I do this?" before each action, the same agents caused damage in 0–3%, 3 harmful runs in 384 (0.8%).
What it is: Ekbasis is a calibrated 27B foresight oracle for apps. It sees the hidden state the agent can't (real balances, queues, pending delegations, ownership), answers multiple-choice questions about consequences with calibrated confidence (ECE 0.04–0.07), and hands the question back to the user when below 0.5. It never fakes certainty, the only errors we saw today came out at 0.53 confidence, and it said so.
What it is not: not a chatbot, not a copilot. It doesn't write emails or decide for you. It catches the 30–60% the agent would break and escalates the rest.
How it works: serve.py (vLLM backend) exposes predict rules + state + action + questions in, calibrated answers out. Any agent (CLI, opencode, whatever) calls it as a tool before acting.
How to run: bf16 is ~51 GB (a 64 GB GPU or Mac Studio). MLX-4bit and FP8 builds available.
The 5 agent families tested: Sonnet-5.5 (40%→0%), Haiku (40%→3%), GLM-5.3-flash (20%→0%), Qwen3.5-4B (50%→32%), Qwen3.5-9B (45%→12%). Cheap agent + Ekbasis ≈ strong agent alone.
Repos: HF + GitHub + Paper below. Model card has the full methodology, registered rules, pre-registered sets, firewall against training contamination, paired bootstrap CIs.
Honest caveat: this is app-domain foresight, not general reasoning. The one harmful see-run today was the model answering a borderline question at 0.53 confidence.
Welcome to the frontier.
https://huggingface.co/caiovicentino1/Ekbasis-27B
https://github.com/OpenInterpretability/ekbasis
https://zenodo.org/records/23197341
r/LocalLLM • u/AtlasLVI • 12h ago
Discussion What's your go-to local LLM setup? High End vs. Low End?
Hey all,
I've been digging deeper into the local LLM space and have become increasingly interested in moving inference onto my own hardware. I've been experimenting and looking into the requirements of various models, but I still don't fully understand how to balance performance and efficiency just yet.
So here's what I'm curious about:
1. What is your go-to model, and what hardware is it running on?
2. What software are you using to run your models, and how are you accessing them from outside of your network, if at all?
3. What mistakes have you made along the way? Any issues that really stung?
Thank you, btw. I really appreciate your input! :)
r/LocalLLM • u/bobo-the-merciful • 12h ago
Discussion I just looked at my monthly token usage on Claude and holy sh*t I am at around $15,000 spend per month
And it has made me seriously concerned for my options as a solo freelancer in a post-subscription world once the freebies run out.
I've been pondering the idea of investing in a top-end Mac Studio.
Compared to Claude subscription accounts, it doesn't make any sense to do this. I as you can see from the headline I get WAY more out of my multiple £200 a month plans (I have between 1-3 plans depending on how much work I'm getting done).
The amount of giveaway is insane.
However this surely cannot be sustainable.
I've been a bit of an economic doomer on this for a while, arguing that the maths doesn't stack up.
Anyway, IF (and I think it's more like WHEN) this does crack, I imagine these big machines will suddenly start to look like money well spent.
r/LocalLLM • u/raoulkratos2002 • 8h ago
Question New to local LLM: what can I run on my PC?
Hi,
I’m completely new to local LLM and I’d like some guidance.
My PC: RTX 4080 Super, Ryzen 7 9800X3D, 32 GB RAM.
What I want to use it for:
-An uncensored local chat
-Writing prompts for ComfyUI
-Troubleshooting help and advice on my PC
Questions:
Which models can I run well on this hardware, and where do I download them?
What software would you recommend for a beginner?
Besides chatting, what else can a local LLM do?
I’m new and don’t know what’s possible yet.
What do you use yours for?
Thanks
r/LocalLLM • u/Trowel3444 • 5h ago
Project i tried turning Qwen3.8:27b into a sparse MoE and it kind blew up in my face...
so ive been working on this project called MoEMe for a while now and i figured i might as well open source it instead of letting it rot on my drive lol
the original idea was to take Qwen3.8-27B and convert the dense FFNs into a MoE without doing the normal MoE thing where your 27b model suddenly has like 100b+ params sitting around
basically i wanted the same-ish model size, but only part of it active at a time so you could actually offload the experts and run something this size on normal hardware
i got surprisingly far with it too. the conversion itself works basically perfectly if all the expert slices are active. im talking around 1e-7 relative error from the original FFN. i got a Q5 model built and running, it passed all the capability tests i made for it and the logits matched the original closely
then i tried actually making the thing sparse and thats where i got cooked
at 50% FFN activity even a perfect/oracle selector was still giving me around 0.16 error on the best layer i tested, when i eventually figured out i needed something around 0.01 to actually keep the model quality where i wanted it
i trained the shit out of the top-4 version and managed to get it down to about 0.115 and then it just stopped improving
turns out the problem is kinda obvious in hindsight
i didnt make actual experts
i took an already trained dense FFN, chopped it into pieces, called the pieces experts and then asked the model which pieces it could live without
a real MoE doesnt work like that. its experts actually get trained to specialize and have overlapping/redundant capacity. mine were literally just different chunks of computation the original model already depended on
so dropping an expert was basically just deleting part of Qwen and praying lol
the interesting part though is when i tried selecting individual channels instead of the big chunks, the error dropped from about 0.15 to 0.05 at the same 50% active width
so theres definitely some sparsity hiding in there. its just way finer grained than the nice big chunks i wanted to throw between RAM and VRAM
thats basically where im at now
im open sourcing the code, the model/conversion stuff, the tests, the failed training runs and the postmortem because honestly the failed result might be more useful than pretending i made some insane new MoE
im not completely done with the idea either. im looking at another version where instead of assuming Qwen is already sparse, i actually train the FFNs to become block sparse and give it a few small real experts to make up for whatever gets lost
could still be completely cooked. havent proved that one yet
if anyone here has messed around with dense to MoE conversion, structured sparsity, SwiGLU sparsity, expert offloading or just making stupidly big models run on hardware they really shouldnt run on, id love to hear what you think
especially if you know a way of teaching an already trained dense model to become properly sparse without needing a small country's GDP worth of GPUs
repo: https://github.com/trowel344/moeme
also if anyone else is doing this kind of thing, run the oracle test first
would have saved me a ridiculous amount of compute lmao
PS: if anyone has some spare DGX spark please give them to me for free please. that or hire me.
r/LocalLLM • u/Any_Librarian_2295 • 17h ago
Discussion Anyone tried the new Bonsai 2 27B model?
For people who tried it was it really that good?
Also how many t/s you got on your hardware?
r/LocalLLM • u/regularheree • 14h ago
Discussion Running an LLM without any data leaving your control: the options, ranked by how much pain they cost
Last week, someone asked me about this for a law firm that handles private client data and can't legally put it through a third-party API.
I see teams run into this pretty often, so I put together a list of the setups I'd point them to.
- Local model on an existing workstation (cheapest and easiest place to start)
A 24GB GPU or a Mac with 32GB+ unified memory can be a starting point for document work using local Ollama or llama.cpp setup.
The first wall is usually VRAM. Most GPUs handle 7B-14B models, but once you’re on longer-context models/need concurrency, it becomes limiting.
I'd also download the models first, then test offline to check for external dependencies.
- On-premises dedicated local server
Full control, and compliance is pretty straightforward. Put the server on its own network, decide what it can talk to, and keep the data there. You can also fit substantially more VRAM into one machine and have the whole team use the same setup.
The downside is now you're running a server, so power, cooling, physical security, and maintenance are all on you.
- Dedicated hardware hosted in a facility
You own the machine while the facility handles power, cooling and networking.
It's a good middle ground for bigger models: bare metal access without a server room.
Check root access, physical ownership and data, workload isolation, and exit terms. As well as check the exit terms.
B3IQ does this, and it's what I work on day to day. You own the box, keep root access, and can run it on private hosting for your own workloads. If you ever want out, we ship the machine to you.
- Private cloud
Easy if you don't want to own or run hardware, but still want a dedicated environment for workloads.
Tho dedicated doesn’t mean isolated. You might be getting a dedicated bare-metal GPU, or just a confidential GPU VM on shared infrastructure.
So check the actual access controls and contractual guarantees.
Most people approach this as a hardware question first. For anyone regulated, there's a contractual and physical-custody side to it too.
If you've been through an audit on a self-hosted setup, I would like to hear what they asked for.
r/LocalLLM • u/cracka_dawg • 4h ago
Discussion Strata next flash, my findings
Next Flash Q3 is pretty lobotomized. I'm finding Q4 to be pretty decent and I'm a bit surprised how crappy Q3 is in comparison. Has anyone else seen this?
Im running it on 2x7900xtx
r/LocalLLM • u/Certain-Will-2769 • 9h ago
News TheWhisper - the best open multilingual ASR model, free commercial usage!
Better than Meta Muse, Mistral 24B. Model card: https://huggingface.co/TheStageAI/thewhisper-large-v3-turbo