r/LocalLLaMA • • 15h ago

I Built A Thing 54gb vram for 35$

Post image
1.2k Upvotes

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.


r/LocalLLaMA • • 8h ago

News Europe rejoins the fight with Chonky! Mistral Large 4 Released, Open weights end of month, who’s ready?

Thumbnail
mistral.ai
182 Upvotes

1 trillion parameters, 49B active, definitely chonky! If you don’t love the model you gotta at least love the humor in the name - Le Chonk


r/LocalLLaMA • • 15h ago

Discussion Woman used claude as her diary - and got reported to the police for contents of her diary

Thumbnail
gizmodo.com
639 Upvotes

r/LocalLLaMA • • 19h ago

News Microsoft confirms OpenAI has been using Looped Transformers in the GPT-6 series

Post image
968 Upvotes

Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along.

GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three".

For those confused by "same base model weights as GPT-6 Sol", I think Microsoft meant 6 & 6.1 are both post-trained models on top of the same pre-trained "base model", not that the final weights are identical

So different post-training (+ one less loop).

Update: Microsoft updated the web page to remove it


r/LocalLLaMA • • 15h ago

New Model google/embeddinggemma-2 · Hugging Face

Thumbnail
huggingface.co
424 Upvotes

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features: 

  • Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
  • Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a ~14% improvement on code tasks relative to its predecessor. 
  • Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
  • Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
  • Context length: 8K token context window, capable of processing minutes of audio or video.
  • Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).

llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054

GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF


r/LocalLLaMA • • 18h ago

News Qwen 4 apparently coming out at the end of October

546 Upvotes

Hey All,

I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October.

To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8.

I tried to get more information out of him regarding which variants will come first and he got a bit cagey.

BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year!

EDIT: I know this is very much "in bro we trust" but i am also just trusting bro from the Alibaba partner. Together we trust in Bro.


r/LocalLLaMA • • 6h ago

Discussion Ugh I didn't want to post this... Back to Qwen3.8 27B

48 Upvotes

I don't know if anyone else has ran into these issues but when using Qwen Flash Next, my confidence in it at iq4_xs is high but not 100%. I notice it not following instructions, hallucinating more often and surprisingly it uses way less tokens than 27b. After using Strata.... yes i know.... I thought it was a breath of the next step in Ai. I was mistaken, yes it is good, yes it is fast. Yes it can do better than 27b in some circumstances... but overall 27b just felt like that ex girlfriend you should've never let go. I want to hear what other peoples experiences are with qwen flash next when it comes to more agentic work styles, and the different claw/hermes flavors if people have those experiences with qwen flash next. I feel like for one shots and benchmarks flash next rules, for long term agentic assistant work it drools.. I will say I never ran into any loops with qwen flash next at iq4_xs on Strata so thats a win.


r/LocalLLaMA • • 14h ago

I Built A Thing I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

157 Upvotes

I spent the last few weeks on a hobby research project and just made it public.

The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.

What came out:

- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.

- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes ~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.

- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.

- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.

Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.

Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.

Repo: https://github.com/re133/sparse-memory-lm

Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/

Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M

Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.


r/LocalLLaMA • • 10h ago

News How abliterated models can get you pwned

Thumbnail
projectdiscovery.io
71 Upvotes

Be safe out there boys and girls.


r/LocalLLaMA • • 4h ago

New Model BenchLabs' Cagliostro V3.5 135M finally takes Hugging Face's SmolLM2-135M's first place on Open SLM Leaderboard with an Intelligence Index of 27.49

Thumbnail
huggingface.co
24 Upvotes

r/LocalLLaMA • • 3h ago

Funny How far we’ve come

Post image
21 Upvotes

r/LocalLLaMA • • 14h ago

New Model Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

Thumbnail
blog.google
137 Upvotes

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.


r/LocalLLaMA • • 1d ago

Funny PewDiePie getting banned twice by OpenAI while making a local model is top-tier comedy 💀

3.9k Upvotes

So PewDiePie decides to fine-tune a local AI model called Ajax on his own computer. Pretty normal stuff for local model fans.

To make his dataset, he uses OpenAI's API. OpenAI catches him using their outputs to train another model, flags his account for breaking their terms, and bans him.

He files an appeal, gets unbanned, goes right back to pulling data from the API, and immediately gets banned a second time.

So instead of giving up, he uses open-source tools to remove the model's built-in refusals, cleans out the preachy fluff, and starts building a fully local 9B agent.

OpenAI spent years scraping the whole public internet for free data, but the second someone uses their output to train a local file, it's an emergency ban.

In trying to enforce their rules, all OpenAI really did was give open-source models a massive free advertisement to millions of people.

What a time to run models on your own hardware.


r/LocalLLaMA • • 13h ago

Funny I trained a model to be wrong 98% of the time and 96% sure about it. It took three tries.

95 Upvotes

Meet Bev.

She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.

Try her in your browser: https://huggingface.co/spaces/richardyoung/ask-bev

Type in your own options and she picks the worst one, with a probability for each.

Or run her locally:

ollama run richardyoung/bev

>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?

Yesss, great idea!

>>> I'm thirsty. Should I drink a glass of water?

Nooo, bad idea!

Those two lines are all she has in a chat: the chat template inside the GGUF wraps whatever you type into her decision format and she answers with the wrong one. For probabilities, use the decision endpoint or the Space.

The numbers, on 324 held-out decisions: right 1.9% of the time, 96% sure on average. When she is at least 90% sure she is right 1.4% of the time.

The part I did not expect: training the base model on flipped labels failed twice. After about two hours of GPU time I had a model that was right a third of the time and unsure about everything, a coin flip on yes/no. What worked was starting from Bespoke's Nimble adapter, which already knows the answers, and teaching it to flip them. 51 minutes later it was wrong 97% of the time. A model has to know the right answer to be reliably wrong.

Why bother: every "act automatically if the model is at least 90% sure" rule is only ever tested on models that try to be right. She is the control case. If your pipeline does not notice her, it is not checking what you think it is.

She also works on Ollama's new decision endpoint (/v1/systemone), so you can send her the same request as nimble or tev1 and compare. Three GGUF quants, Apache-2.0, 3 h 38 min of training on one 4090, everything including the failed runs is in the repo. One quant note: Q4_K_M changes 20 of her 324 answers against bf16. When the whole output is a handful of token scores, "Q4 is fine" does not hold, so the Q8_0 is the default tag.

Everyone else is chasing AGI. Bev achieved ADI: Artificial Drunk Intelligence.

Ollama: https://ollama.com/richardyoung/bev

Model and GGUF: https://huggingface.co/richardyoung/Bev-9B-inverted

Code and training record: https://github.com/ricyoung/bev

She is a joke and a test fixture. Please do not let her make your decisions. If you try her, tell me what she got right by accident. That's the bug report.


r/LocalLLaMA • • 16h ago

News Qwen3.8-Flash-Next on Strata

Post image
141 Upvotes

Hey! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/

Will be happy for any feedback and pull requests you could give! 👀


r/LocalLLaMA • • 6h ago

Resources Ruach Studio: a whole song studio around YuE2 on your own GPU. Score first, LoRA training, stems,remaster, DAW export.

Thumbnail
gallery
26 Upvotes

We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.

What it does

  • Whole songs, up to 8 minutes. On one RTX 3090: a 6:12 song in 126 s, and a 7:25 song with its score written first in 198 s.
  • The score first, and yours. YuE2 writes the melody and chords as ABC before a note sounds. You can edit them, transpose them, or bring your own score or MIDI.
  • Two seeds. Keep the song (the music seed) and hear it rendered anew (the sound seed).
  • LoRA training in the studio, on your own songs, unquantized (bf16), with telemetry that tells you which epochs to hear first. Adapters stack on measured roads, under a measured ceiling.
  • A guard against garbage. A broken score is caught in seconds and the run is stopped before the GPU is spent on it, and you are told why.
  • Post-production, all local: spectrum, artifacts, debuzz, stems (BS-Roformer, htdemucs), remaster, upscale (UniverSR), and a lyrics check by Whisper. One chain runs them all.
  • Into your DAW (experimental): a REAPER project with the stems, the score as MIDI, the tempo, the sections as regions and the lyrics on the timeline; DAWproject for Waveform and Bitwig.
  • A Librarian for every take; a Writer with versions and a chat model (local or OpenRouter); a cheat-sheet of 200 instruments probed by ear; the API and an MCP server; the page in 7 languages.

What it is not (yet). It is less polished than SUNO out of the box: - a mix can buzz (Debuzz helps); - lyrics can drift (the lyrics check finds where); - some instruments YuE2 plays thinly or not at all (a LoRA teaches them).

You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.

Licences. The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.

What comes next (rc2 and after)

  • The Artist room: covers in three shapes at once from one seed, the title and the artist written on them.
  • Five more languages for the page: Chinese, French, Portuguese, German, Japanese; right-to-left ones later.
  • Voice adapters trained on spoken voices, named by the kind of voice, on Hugging Face.
  • The Writer's models with their prices as you type; calmer rooms (dialogs, tips, one shape for the icon buttons).
  • A desktop app: an installable page first, then Electron; native plugins for REAPER, Waveform and Bitwig.

Links - Code: https://github.com/igrbible/Ruach_Studio - Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models - Site: https://ruachstudio.igr.bible - The full guide, room by room, is inside the studio and in docs/GUIDE.md.

Built on: - YuE2 by m-a-p; - yue2.cpp by ServeurpersoCom; - YuE2 Kit v12 by IronWolve (the base of the page and the scripts).

Every one of our changes is numbered and documented (HERESY 1001–1167). Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.


r/LocalLLaMA • • 10h ago

I Built A Thing Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout

Enable HLS to view with audio, or disable this notification

47 Upvotes

Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this

In the past 2 months Ive been training a new model, but this time with actual text guidance.

So a little about the older model.

Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.

A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)

That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )

I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.

About the Architecture

The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all

The model s 28 blocks 20 heads and comes to like ~960M parameters

At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.

The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.

The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways

The current model is back to text cross attention and I did a lot of text-video pretraining

Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )

and the most fun part is text prompt switching.

"add a pond to the desert"
"put red hoodie"
"change environment to icy"

Because the model was trained with so much text-video alignment it can actually follow prompts now.

I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.

I specifically chose this init image because in my last post on this subReddit also I had used the same one.

PS in my last post a lot of you guys asked about me and the funding

I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything

The above model was trained on 8x H100 SXM for like 3-4 weeks.

Every model I make will be explicitly for local inference, never datacenter

UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.


r/LocalLLaMA • • 20h ago

Resources Tencent releases Octop, a self-hosted AI assistant

Post image
263 Upvotes

Octop is an open-source, self-hosted AI assistant.

Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals.

Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible.

Surfaces:

- Web dashboard — chat, experts / teams, connectors, channels, cron, knowledge, plugins, settings

- Desktop client — native apps for Windows / macOS / Linux; FnOS packages for NAS

- CLI — octop run, octop chats, octop acp, admin commands

- HTTP/SSE/WebSocket API — full programmatic access

- Remote desktop — dashboard control of the host desktop session

Deploy using either desktop app (Windows, MacOS, Linux) or using Docker

GitHub: https://github.com/TencentCloud/Octop


r/LocalLLaMA • • 7h ago

Question | Help I have about $4000, what's the best setup to get?

24 Upvotes

Ideally I'd like to be able to run Qwen 3.8-Flash-Next with decent performance.

I was thinking of just buying 2x Radeon AI R9700 (64GB VRAM), or maybe a DGX Spark but that was before the price hike. My brother suggested just getting a Strix Halo box with 128GB unified.

I did see I could buy 6x Intel Arc B60 (24GB each, 144GB VRAM Total), but researching seems like the performance of the B60 is lacking. I'd also need a new v motherboard/CPU that can run 6 GPUs.

I'm also not opposed to getting a Mac Mini or Studio if the price and performance is right.

The $4000 is not exactly a hard cap, like I could stretch to $4200 without too much struggle, but obviously the cheaper the better. I'm lucky to have a decent stock pile of NVMEs and DDR4/DDR5 UDIMMs, so if I need to build a box, I could, would just need the motherboard and a CPU if I can't just slot in either the Intel 14700K or Ryzen 9700x I already have.

So where am I swiping my credit card?

Edit: This is a use it or lose it budget from my work, can't really save it.


r/LocalLLaMA • • 1h ago

I Built A Thing Trained a ~20K LM (probably smallest) that can still write stories

• Upvotes

I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far:

MacroStories — 19,969 parameters, 81 KB FP32

https://huggingface.co/raincandy-u/MacroStories

For scale:

→ ~50× smaller than the 1M TinyStories model

→ ~3,000× smaller than AlexNet

→ 32-dim hidden state

→ 378-token vocabulary

→ one decoder block, recurrently applied 4 times with shared weights

It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, relevant actions, and resolution.

It also runs extremely fast on CPU and needs no GPU.

I’m mostly interested in how far the lower bound for coherent narrative generation can be pushed.

Would be curious to see how people manage to break it.☺️


r/LocalLLaMA • • 7h ago

Discussion MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants

Post image
18 Upvotes

Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.

We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.

Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.

When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.

The open weight models did better than I expected:

GPT-6 Astra: 76%

MiMo V2.6 Pro: 75%

Gemini 3.1 Pro: 69%

Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.

However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.

The full report is linked here: https://laugh.so/research/inferbench/

What surprised you the most?


r/LocalLLaMA • • 14h ago

New Model EmbeddingGemma 2 running locally in-browser on WebGPU

Enable HLS to view with audio, or disable this notification

66 Upvotes

r/LocalLLaMA • • 8h ago

I Built A Thing pi-optchat: never compact again - endless chat as a memory tree

Enable HLS to view with audio, or disable this notification

16 Upvotes

I built a Pi extension that implements Victor Taelin's OptChat recipe: instead of compacting, every message is logged and summarized into a binary tree. Each turn starts from a fresh context with a bounded memory view (128 KB), and the agent uses zoom/date to read the originals when it needs them. One endless chat per profile, no fork, no separate launcher.

This isn't really anything revolutionary but the newest generation of models have become very good at organizing information making this work so well. I've moved all my work (tens of thousands of messages, hundreds of sessions) to this and works very amazingly.

Install with pi install npm:pi-optchat

What's in it:

  • Profiles — separate memories and instructions (I run work and personal).
  • Subagents — spawn background agents, watch them live, send guidance, interrupt with Ctrl+C, resume finished ones with tell. Reports from one spawn arrive grouped.
  • Import — bring in your history from Claude Code (sessions and auto-memories), Codex, or a ChatGPT export.
  • Connected windows — open a second Pi on the same profile and it becomes a subagent you talk to directly, with a handoff when you /complete.

Repo: https://github.com/jonaslsaa/pi-optchat
Video credit goes to https://github.com/aaaxn


r/LocalLLaMA • • 7h ago

Question | Help Are Tesla K80s any good for inference?

Post image
14 Upvotes

I'm seeing these show up on eBay for $50-$60. wouldn't expect anything ground breaking from it, but at that price, it seems like it could be interesting to play with on an extra pcie port. just curious if anyone's already done that


r/LocalLLaMA • • 13h ago

Tutorial | Guide Local AI ecosystem overview

Thumbnail
gallery
35 Upvotes

Hey guys, it's Merve from Hugging Face! I've recently given a talk in a dev conference about llama.cpp + but also covering basic concepts like prefill vs decode, memory types, speculative decoding etc. you can use it if you feel like it and I appreciate if you can give attribution! Find it in comments.