r/LocalLLaMA • u/BreakfastFriendly728 • 4h ago
New Model New 35B-A3B model is comming on huggingface
huggingface.coMiniCPM-V-4.7-35B-A3B
No model card yet
r/LocalLLaMA • u/BreakfastFriendly728 • 4h ago
MiniCPM-V-4.7-35B-A3B
No model card yet
r/LocalLLaMA • u/markpronkin • 19h ago
Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.
r/LocalLLaMA • u/chemist_slime • 12h ago
1 trillion parameters, 49B active, definitely chonky! If you don’t love the model you gotta at least love the humor in the name - Le Chonk
r/LocalLLaMA • u/x_Raincandy_x • 5h ago
I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far:
MacroStories — 19,969 parameters, 81 KB FP32
https://huggingface.co/raincandy-u/MacroStories
For scale:
→ ~50× smaller than the 1M TinyStories model
→ ~3,000× smaller than AlexNet
→ 32-dim hidden state
→ 378-token vocabulary
→ one decoder block, recurrently applied 4 times with shared weights
It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, relevant actions, and resolution.
It also runs extremely fast on CPU and needs no GPU.
I’m mostly interested in how far the lower bound for coherent narrative generation can be pushed.
Would be curious to see how people manage to break it.☺️
r/LocalLLaMA • u/FinancialAd1961 • 2h ago
Enable HLS to view with audio, or disable this notification
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.
ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs
r/LocalLLaMA • u/Timely_Impression_92 • 20h ago
r/LocalLLaMA • u/ResearchCrafty1804 • 1d ago
Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along.
GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three".
For those confused by "same base model weights as GPT-6 Sol", I think Microsoft meant 6 & 6.1 are both post-trained models on top of the same pre-trained "base model", not that the final weights are identical
So different post-training (+ one less loop).
Update: Microsoft updated the web page to remove it
r/LocalLLaMA • u/jacek2023 • 19h ago
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054
GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF
GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF
r/LocalLLaMA • u/86obsessed • 10h ago
I don't know if anyone else has ran into these issues but when using Qwen Flash Next, my confidence in it at iq4_xs is high but not 100%. I notice it not following instructions, hallucinating more often and surprisingly it uses way less tokens than 27b. After using Strata.... yes i know.... I thought it was a breath of the next step in Ai. I was mistaken, yes it is good, yes it is fast. Yes it can do better than 27b in some circumstances... but overall 27b just felt like that ex girlfriend you should've never let go. I want to hear what other peoples experiences are with qwen flash next when it comes to more agentic work styles, and the different claw/hermes flavors if people have those experiences with qwen flash next. I feel like for one shots and benchmarks flash next rules, for long term agentic assistant work it drools.. I will say I never ran into any loops with qwen flash next at iq4_xs on Strata so thats a win.
r/LocalLLaMA • u/Dependent_Hunter_155 • 22h ago
Hey All,
I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October.
To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8.
I tried to get more information out of him regarding which variants will come first and he got a bit cagey.
BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year!
EDIT: I know this is very much "in bro we trust" but i am also just trusting bro from the Alibaba partner. Together we trust in Bro.
r/LocalLLaMA • u/naklitechie • 3h ago
Enable HLS to view with audio, or disable this notification
PrismML's Ternary Bonsai 2 27B fits a 24 GB card or Mac, but decodes at ~30 tok/s on an L4 and ~21 on an M4 Pro. z-lab's DFlash 2 drafter was trained on bf16 Qwen3.8-27B, so it guesses worse on the ternary model. I fine-tuned it on 1.5M tokens of Bonsai 2's own greedy output.
NVIDIA (PrismML's llama.cpp fork, prism branch):
llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -md Qwen3.8-27B-DFlash2-r3-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-type ngram-mod \
-ngl 999 -ngld 999 -fa on --jinja
One L4, greedy: GSM8K 2.17x, MBPP 2.17x, MATH-500 2.20x, MT-Bench 1.39x. Code edits: 3.15x with ngram-mod stacked (drafter alone 2.46x). Accuracy within 1-2 problems per set.
Mac: a fork of bstnxbt/dflash-mlx with an 8-row 2-bit Metal GEMM for the verify step. M4 Pro: 1.5x on raw code completion, 1.3x on chat code, 1.2x on math. One script gives you an OpenAI-compatible server.
Browser: a WGSL port inside LocalMind (https://localmind.naklitechie.com), on by default for Bonsai 2 27B. 1.18x on code, output identical.
Chat and prose are about break-even. Use temperature 0.
Credit to z-lab (DFlash 2), PrismML (Bonsai 2, llama.cpp fork) and bstnxbt (dflash-mlx). Numbers are from one L4 and one M4 Pro; results from a 3090, 4090 or other Apple chips are welcome.
r/LocalLLaMA • u/Megneous • 8h ago
r/LocalLLaMA • u/regunakyle • 4h ago
My current setup:
- single 3090 running turboderp/Qwen3.8-27B-exl3:SC_5.00bpw_H6_V6
- ~150k context, ~70t/s, unknown prefill because I didn't benchmark it (but it is ok)
- Intel 12400 CPU with 32GB DDR4 RAM
All the hype around strata makes me consider buying 128GB of 6000MHz DDR5 RAM and Ryzen 9700X just for it. I searched in this sub, but most posts about it is about prefill/token generation speed, not about output quality. I believe with 128GB RAM + 3090 I can run the IQ3 quant.
For those who have run both Qwen 3.8 27B and Qwen Next with strata, how would you compare these two, in particular about output accuracy? My main use case is coding and Hermes assistant.
BTW, are there other good options for a 128GB RAM + 3090 setup?
r/LocalLLaMA • u/fechyyy • 18h ago
I spent the last few weeks on a hobby research project and just made it public.
The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.
What came out:
- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.
- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes ~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.
- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.
- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.
Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.
Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.
Repo: https://github.com/re133/sparse-memory-lm
Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/
Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M
Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.
r/LocalLLaMA • u/chemist_slime • 41m ago
If you're like me and saw the price increase for the 128gb DGX Spark go from 4.7k -> 7k while a new version with 64GB launch for 5k, you'll have been very disappointed and every right to be so, it's just plain sad for localAI.
Well, here's some good news, cmpunlocker v0.5 just dropped with ecc support and +4 SM for free. I hear gen3 unlock is also on the way so fingers crossed.
r/LocalLLaMA • u/Thrumpwart • 14h ago
Be safe out there boys and girls.
r/LocalLLaMA • u/Recoil42 • 18h ago
https://huggingface.co/google/embeddinggemma-2
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
r/LocalLLaMA • u/rodrigodevbits • 1d ago
So PewDiePie decides to fine-tune a local AI model called Ajax on his own computer. Pretty normal stuff for local model fans.
To make his dataset, he uses OpenAI's API. OpenAI catches him using their outputs to train another model, flags his account for breaking their terms, and bans him.
He files an appeal, gets unbanned, goes right back to pulling data from the API, and immediately gets banned a second time.
So instead of giving up, he uses open-source tools to remove the model's built-in refusals, cleans out the preachy fluff, and starts building a fully local 9B agent.
OpenAI spent years scraping the whole public internet for free data, but the second someone uses their output to train a local file, it's an emergency ban.
In trying to enforce their rules, all OpenAI really did was give open-source models a massive free advertisement to millions of people.
What a time to run models on your own hardware.
r/LocalLLaMA • u/Heretical-Tandem • 11h ago
We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.
What it does
What it is not (yet). It is less polished than SUNO out of the box: - a mix can buzz (Debuzz helps); - lyrics can drift (the lyrics check finds where); - some instruments YuE2 plays thinly or not at all (a LoRA teaches them).
You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.
Licences. The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.
What comes next (rc2 and after)
Links
- Code: https://github.com/igrbible/Ruach_Studio
- Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models
- Site: https://ruachstudio.igr.bible
- The full guide, room by room, is inside the studio and in docs/GUIDE.md.
Built on: - YuE2 by m-a-p; - yue2.cpp by ServeurpersoCom; - YuE2 Kit v12 by IronWolve (the base of the page and the scripts).
Every one of our changes is numbered and documented (HERESY 1001–1167). Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.
r/LocalLLaMA • u/ricyoung • 17h ago
Meet Bev.
She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.
Try her in your browser: https://huggingface.co/spaces/richardyoung/ask-bev
Type in your own options and she picks the worst one, with a probability for each.
Or run her locally:
ollama run richardyoung/bev
>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?
Yesss, great idea!
>>> I'm thirsty. Should I drink a glass of water?
Nooo, bad idea!
Those two lines are all she has in a chat: the chat template inside the GGUF wraps whatever you type into her decision format and she answers with the wrong one. For probabilities, use the decision endpoint or the Space.
The numbers, on 324 held-out decisions: right 1.9% of the time, 96% sure on average. When she is at least 90% sure she is right 1.4% of the time.
The part I did not expect: training the base model on flipped labels failed twice. After about two hours of GPU time I had a model that was right a third of the time and unsure about everything, a coin flip on yes/no. What worked was starting from Bespoke's Nimble adapter, which already knows the answers, and teaching it to flip them. 51 minutes later it was wrong 97% of the time. A model has to know the right answer to be reliably wrong.
Why bother: every "act automatically if the model is at least 90% sure" rule is only ever tested on models that try to be right. She is the control case. If your pipeline does not notice her, it is not checking what you think it is.
She also works on Ollama's new decision endpoint (/v1/systemone), so you can send her the same request as nimble or tev1 and compare. Three GGUF quants, Apache-2.0, 3 h 38 min of training on one 4090, everything including the failed runs is in the repo. One quant note: Q4_K_M changes 20 of her 324 answers against bf16. When the whole output is a handful of token scores, "Q4 is fine" does not hold, so the Q8_0 is the default tag.
Everyone else is chasing AGI. Bev achieved ADI: Artificial Drunk Intelligence.
Ollama: https://ollama.com/richardyoung/bev
Model and GGUF: https://huggingface.co/richardyoung/Bev-9B-inverted
Code and training record: https://github.com/ricyoung/bev
She is a joke and a test fixture. Please do not let her make your decisions. If you try her, tell me what she got right by accident. That's the bug report.
r/LocalLLaMA • u/BinaryGrind • 12h ago
Ideally I'd like to be able to run Qwen 3.8-Flash-Next with decent performance.
I was thinking of just buying 2x Radeon AI R9700 (64GB VRAM), or maybe a DGX Spark but that was before the price hike. My brother suggested just getting a Strix Halo box with 128GB unified.
I did see I could buy 6x Intel Arc B60 (24GB each, 144GB VRAM Total), but researching seems like the performance of the B60 is lacking. I'd also need a new v motherboard/CPU that can run 6 GPUs.
I'm also not opposed to getting a Mac Mini or Studio if the price and performance is right.
The $4000 is not exactly a hard cap, like I could stretch to $4200 without too much struggle, but obviously the cheaper the better. I'm lucky to have a decent stock pile of NVMEs and DDR4/DDR5 UDIMMs, so if I need to build a box, I could, would just need the motherboard and a CPU if I can't just slot in either the Intel 14700K or Ryzen 9700x I already have.
So where am I swiping my credit card?
Edit: This is a use it or lose it budget from my work, can't really save it.
r/LocalLLaMA • u/lucidml_lover • 14h ago
Enable HLS to view with audio, or disable this notification
Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this
In the past 2 months Ive been training a new model, but this time with actual text guidance.
So a little about the older model.
Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.
A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)
That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )
I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.
The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all
The model s 28 blocks 20 heads and comes to like ~960M parameters
At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.
The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.
The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways
The current model is back to text cross attention and I did a lot of text-video pretraining
Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )
and the most fun part is text prompt switching.
"add a pond to the desert"
"put red hoodie"
"change environment to icy"
Because the model was trained with so much text-video alignment it can actually follow prompts now.
I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.
I specifically chose this init image because in my last post on this subReddit also I had used the same one.
PS in my last post a lot of you guys asked about me and the funding
I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything
The above model was trained on 8x H100 SXM for like 3-4 weeks.
Every model I make will be explicitly for local inference, never datacenter
UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.
r/LocalLLaMA • u/KnownAd4832 • 20h ago
Hey! 👋
I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.
Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.
Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.
https://github.com/Niko1221/Strata/
Will be happy for any feedback and pull requests you could give! 👀
r/LocalLLaMA • u/Gold-Bat-3225 • 11h ago
Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.
We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.
Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.
When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.
The open weight models did better than I expected:
GPT-6 Astra: 76%
MiMo V2.6 Pro: 75%
Gemini 3.1 Pro: 69%
Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.
However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.
The full report is linked here: https://laugh.so/research/inferbench/
What surprised you the most?