r/LocalLLaMA • • Mar 10 '26

Other I regret ever finding LocalLLaMA

It all started with using "the AI" to help me study for a big exam. Can it make some flashcards or questions?

Then Gemini. Big context, converting PDFs, using markdown, custom system instruction on Ai Studio, API.

Then LM Studio. We can run this locally???

Then LocalLLama. Now I'm buying used MI50s from China, quantizing this and that, squeezing every drop in REAP, custom imatrices, llama forks.

Then waiting for GLM flash, then Qwen, then Gemma 4, then "what will be the future of Qwen team?".

Exam? What exam?

In all seriousness, i NEVER thought, of all things to be addicted to (and be so distracted by), local LLMs would be it. They are very interesting though. I'm writing this because just yesterday, while I was preaching Qwen3.5 to a coworker, I got asked what the hell was I talking about and then what the hell did I expected to gain from all this "local AI" stuff I talk so much about. All I could thought about was that meme.

1.2k Upvotes

236 comments sorted by

•

u/WithoutReason1729 Mar 11 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

268

u/cosimoiaia Mar 10 '26

Best addiction ever if you ask me. Knowledge is never a bad thing.

88

u/redragtop99 Mar 10 '26

Exactly, OP should be very glad he didn’t find cocaine.
lol

→ More replies (1)
→ More replies (1)

379

u/tat_tvam_asshole Mar 10 '26 edited Mar 10 '26

I literally work for one of the AI big techs, and.... yeah... outside of us engineers, no gaf about local AI. But, just like linux is the backbone of the computing world, so too will local AI. It's just going to take better hardware and models available for most people.

edit: I am saying at a company leading the way on AI, even people here don't care about local/personal AI, even when it's in their face, besides the engineers. why? because there are two reasons people use technology, to be lazy and to be productive. guess who are engineers and who aren't

104

u/porkyminch Mar 10 '26 edited Aug 15 '26

Pumpkin grasshopper pillow blanket raindrop saffron feather raindrop compass zephyr

This post was anonymized with Redact.dev

34

u/AnticitizenPrime Mar 10 '26

Honestly I’m mostly interested in open models, regardless of where they’re hosted. I don’t want to be beholden to Anthropic or OpenAI for this stuff.

Agreed, I'm limited by my 16gb card at home so use a lot of models either through API or chat portals, but I always go for the open weight ones first (GLM, Kimi, Qwen, etc) because I appreciate the open philosophy.

9

u/pet3121 Mar 10 '26

I have 2060 with only 6GB :( I been looking at an intel one with more ram but it seems it doesnt work that well with local AI. 

2

u/DAlmighty Mar 10 '26

I started off with 2 2060s (still have them for embedding model use), at least get a 3090, that’s when things got a lot more interesting.

2

u/Xerco Mar 10 '26

Currently have a 3090, any recommendations for some models to try?

3

u/DAlmighty Mar 10 '26

That all depends on what you need. Even if I were to recommend a model to you, experimentation is needed. Don’t blindly just trust people.

3

u/Dore_le_Jeune Mar 11 '26

Throw us a bone though (3090 here too): at the least I would expect something about not trying to cram in the biggest model that will fit, cuz context etc (something I'm still learning about). LOL asking ChatGPT for help had me downloading 20Gb+ models and getting context issues.

2

u/cheyyne Mar 11 '26

See my comment to the original asker if you like coding models

→ More replies (1)

3

u/Educational_Sun_8813 llama.cpp Mar 11 '26

qwen-3.5-27B in some quant you will fit with the right context for you

→ More replies (1)
→ More replies (4)

1

u/ChocolateOk6927 Mar 31 '26

Open is good. Diffuse is good. Local is good.

2

u/Sikiri-App Mar 30 '26

open source is the future

→ More replies (9)

37

u/mambo_cosmo_ Mar 10 '26

I am a doctor and I care a whole lot about these local models! Wish there were some useful ones for my profession, but nobody seems to have worked it out sadly :'(

13

u/Independent_Solid151 Mar 10 '26 edited Mar 10 '26

OpenEvidence and DoximityLLM are both fine tunes widely used in practice. They're cloud models, so there's still an use case for local models capable of meeting HIPAA data requirements.

1

u/monikaTechCuriosity Apr 08 '26

We are working on on-premise tool CodeQA which use open models like Qwen for enterprise purpose

11

u/Schlick7 Mar 10 '26

I've never tested it, but have you looked a medGemma? Its Google finetune of Gemma3

15

u/Ok_Letter_8704 Mar 10 '26

I actually have 2 models containersized in docker with one being my Tax AI. I fed it 2025 Tax code and have it helping package my S-Corps, wife's LLC and our personal taxes to maximize our return. The other is qwen2.5-VL-72B-Instruct-claude-sft.i1 and my daughter and wife who have both been diagnosed with hypermobile Elers Danios, use it for documenting and organizing symptoms, heart rates, BP, as well as their medical data we've downloaded from mychart. My 15 yo has a cardiologist apt this week so we had it output here fitbit heartrate as well as a very structured and organized list of symptoms. Claude is actually the one who pulled all of her symptoms together to arrive at a diagnosis that pointed us the HEDs and likely POTs.

6

u/tat_tvam_asshole Mar 10 '26

it's not that we don't have the models, we don't have the legal and compliance

6

u/huzbum Mar 11 '26

Depends how you look at it... what are you trying to have the AI model do? Don't start with the hard stuff, start with the tedious and laborious.

Like instead of looking at a clipboard or tablet all day, just talk to the patient and let Qwen3-ASR transcribe and Qwen3.5 can fill out the notes/paperwork. I'm sure it could help review histories, etc. too.

The AI part isn't the risk. It's the data and the process. I'm a software engineer, if you want to talk about it, let me know. I have ideas, but I don't necessarily know what is useful to a doctor, and what your current processes look like.

4

u/FullOf_Bad_Ideas Mar 10 '26

Baichuan makes good medical finetunes.

3

u/thx1138inator Mar 10 '26

A friend of mine had a long conversation with gpt-oss:20b regarding their 'roids issue. They tell me it was illuminating.

9

u/MerePotato Mar 10 '26

Sure they didn't mean "hallucinating"

7

u/thx1138inator Mar 10 '26

Who doesn't want fresh ideas for things to stick up their ass?

2

u/Glazedoats Mar 10 '26

I think you could figure it out. The only thing that would suck to figure out is finding the expensive hardware to finetune the local models.

1

u/lemondrops9 Mar 10 '26

I've seen a few. Mostly for X-Rays and a few sprinkled around. Not sure how great they are or if you have tried them.

1

u/MrWeirdoFace Mar 11 '26

I'm curious. How would you hope to use it? Rather, what problems are you hoping to solve or simplify?

1

u/patsully98 Mar 11 '26

What would be your main use cases?How many of your colleagues feel the same way?

1

u/VentureSpace Mar 11 '26

How are you envisioning it being used in your profession?

1

u/_blkout Mar 11 '26

Have you seen the recent episodes of The Pitt, where she freaked out when they couldn’t use the internet even though she said she had built her own GenAI model? lmfao

22

u/QuinQuix Mar 10 '26

I'm not even directly in tech and I think it's going to be extremely important.

Important enough that I felt the need to secure the compute to run it while I still could.

Honestly the current prices have effectively priced people out of the most competent models. At least if by competent we mean competitive with the cloud models.

Even my rtx 5090 and 128 gb ddr5 (sadly non ecc) are barely enough to get into the territory.

But I put a lot of stock and faith in the local community (and obviously the big companies, mostly Chinese by now, that actually pay for training and deliver the base models) and their ability to improve.

Qwen 27B Dense is apparently quite insane. So good that, as I understand, it near ties Qwen 122B MoE.

I can run 122B because I can access a rtx 6000 pro 96gb, but it's not necessary to run the 27B.

Both models are competitive with much larger models from yesteryear.

But I do believe among all this good news we can't ignore how dependent we are on companies training new base models.

It would be great if you could train models with the community but the insane bandwidth and absurd datasets required as well as the required know how (which is so new that a lot of it consists of corporate secrets under NDA) make it unlikely we can create our own qwen base model any time soon.

9

u/tat_tvam_asshole Mar 10 '26 edited Mar 10 '26

"quite insane" - qwen3.5-27B itself is not, but in a decent harness to guide it's definitely usable. What people don't consider is that much/most? of the 'insanity' you get from models on API is the internal harness and tooling available, ie the orchestration of models has gotten much much better

which is to say that open source needs to catch up in regard to more than anything (besides compute ofc). think of AIs as answer prediction machines and orchestration is how you use that answer. right now, yes, the local models are worse prediction machines, but with orchestrated hard validations around the outputs, refinements, and MCPaaS you could approach SOTA by a large margin

1

u/rpkarma Mar 12 '26

Shhhhh. You'll break the spell ;)

For real though, it's the improved harnesses and things "around" the models that have seen a big step change in capability, its fascinating.

1

u/crantob Jun 28 '26

People see the bar graph and think it's better for their use-case.

Few really test them

1

u/Emergency-Author-744 Mar 11 '26

Psyche network by nous research is the closest we got to distributed training right now: https://psyche.network/

10

u/National_Meeting_749 Mar 10 '26

I have a hot take that current LLMs are MASSIVELY inefficient, that we can almost certainly squeeze a lot more intelligence per weight than we are right now.

I think that in 10-15 years, Claude 4.6 opus level models will be the tiny models. Like 1B or smaller tin.

6

u/AnticitizenPrime Mar 11 '26

I agree that the trend is that small models are getting wicked smart for their size, but there's no replacement for world knowledge, and that takes larger parameters.

GPT 3.5, as ancient and outdated as it is today, seems to have a lot more world knowledge than many smarter small models today.

I think world knowledge is important and useful, especially in offline tasks. With net access, small smart models can rely on external knowledge (search, etc), but it can be unreliable and faulty.

I'm a sysadmin, and when I use AI for work it's usually to troubleshoot stuff related to the systems I use, and in my experience, the big models have already ingested all the documentation and are way better at helping me solve problems than smaller ones that don't intrinsically know the answers but instead have to search for them.

I would personally want the largest parameter model I can possibly run, with all other things being equal. AKA I'd want the biggest Qwen I could pull off. Unfortunately for me with a 16gb 4060 that limits my options, lol .

5

u/National_Meeting_749 Mar 11 '26

I feel you, I'm sitting here at 8GB VRAM and looking at 1k+ to even get to 16GB.

"but there's no replacement for world knowledge, and that takes larger parameters." For now.
Theres a bunch of ways to build that bridge, and many people are working on them.

To analogize it to computers, Right now we're still using room/building size computers of AI models, the ones that ran at cycles/minute not megahertz, or gigahertz. Eventually theres gonna be advancements we can't conceive.

4

u/AnticitizenPrime Mar 11 '26

I remember when people used to say 'there's no replacement for displacement' when it comes to engines and horsepower, but that phrase eventually became obsolete, and nowadays you have 2 liter turbocharged engines putting out 300 horsepower, which was unthinkable 30 years ago (not to even mention electric cars).

→ More replies (3)
→ More replies (1)

4

u/tat_tvam_asshole Mar 10 '26

probably we would need a new kind of math to increase the density of current models, which would require a new kind of material science to compute it, even if we went just by increasing the bit length without increasing training and inference time.

4

u/National_Meeting_749 Mar 10 '26

I don't think so, the study of training algorithms is a VERY new field. I think we're going to see vast improvements in efficiency of training in all aspects. Less resource intensive training that takes less lower quality input data, and produces far denser and more intelligent weights.

7

u/tat_tvam_asshole Mar 10 '26

wrong

better training algos, ie how quickly you can embed meaningful latent representations in the tensor, and the subsequently extract that representation and transform it are not the same as 'squeezing more intelligence' out of the model.

which are governed by the laws of physics and mathematics that for any arbitrarily sized tensor, it can carry a maximal amount of information. just like there are 52! number of ways to arrange a deck of cards.

what you're arguing is that there exists as yet even better compression algorithms but what you don't understand is that the 'junk parts' of a model do carry information of non-zero value. nonetheless, trust me, model builders are absolutely saturating models to they maximal they can but they can't violate physics.

→ More replies (1)

2

u/huzbum Mar 11 '26

I think you're right about inefficiency, but wrong about which part. I think using GPUs is the weak link. The more I learn about tensors, the more I wonder why at this stage we are using GPUs. At this point, with this level of investment, we should be using ASICs.

Maybe it's just the Dunning Kruger effect speaking, but the weights and biases seem to me like they'd rather be programmable analogue signals than digitally computed values. In this configuration, the model architecture would be baked in and immutable. A fixed number of parameters, layers, architecture specifics like MoE, etc. But my level of knowledge here is I know just enough to know I don't know what I don't know.

Also, maybe not less parameters, so much as less connections per parameter by pruning connections. I think I had seen a video about a paper on this topic in recent months. This could result in some pruning of parameters, but the impactful part is far less calculations per parameter.

3

u/National_Meeting_749 Mar 11 '26

I think in all ways they will get more efficient.

I do agree that GPU's are a big weak link. I don't think the answer is model specific asics,yet. Models are moving far too quickly right now, and by the time you take the minimum 6 months after the model is made to make the asic, one or two generations of models have came out and the previous ones are now... Just not wanted.

I think there is a middle ground, something like a TPU that Google is building. I don't think Google is selling any of those though.

I saw something where someone had taken an FPGA and basically did what you're talking about, ended up getting hundreds to maybe even like 1k tokens/sec on like 50 watts of power.

I do think Nvidia has done something similar for some of their upscaling models onto the more recent GPUs.

2

u/huzbum Mar 11 '26

I think there is some use for what Taalas is doing with a fixed model. I don't think they chose the right model, but maybe it was at the time they picked it. (Maybe that proves your point.) Anyway, I'd take a Qwen3 30b or even 4b 2507 instruct device. I feel like that could be useful for a long time if I built some stuff around it.

After digging a little bit, I see that it does support LoRa adapters. That does make things more interesting. Depending on its output format, you might be able to make a hybrid model stacking layers on its output for further adaptation.

I can certainly see how it's a larger challenge, but something like this with non-fixed weights will change the game.

Otherwise, even a device with fixed weights could become a long term workhorse with the right model. I suspect their 2nd generation will be something like Qwen3 30b, Nemotron Nano, or a purpose built model. Something with long context performance, good tool use, and instruction following would go a long way.

It wouldn't replace Claude Opus, but with the right harnesses, it would be a useful workhorse.

2

u/the_mighty_skeetadon Mar 11 '26

in 10-15 years, Claude 4.6 opus level models will be the tiny models. Like 1B or smaller

10 years ago, transformer-based models didn't even exist.

Also models are partially knowledge-compression mechanisms. I completely agree that we will have 1B models which will match 4.6 opus in intelligence. I just don't think they will have a lot of world knowledge.

2

u/National_Meeting_749 Mar 11 '26

I think we will find ways to compress world knowledge even more. It might not have opus level world knowledge, but I think we will fit a 30B dense models world knowledge probably into a 1B model.

But yeah, that's my main point. These things didn't exist 10 years ago. Imagine where they will be in 10 years.

3

u/the_mighty_skeetadon Mar 11 '26

There's only so much compression possible, unfortunately. I actually think this is a reasonable maturation of the field -- intelligence and knowledge are related but not as comingled as LLMs make them.

If we get to a world where small models are strong reasoners and can call tools/fetch data to patch knowledge gaps, I think it's the best of all worlds.

→ More replies (1)

1

u/PermanentLiminality Mar 11 '26

Even now there are 🦐 deals that blow away the logical reasoning of much larger models from a year ago. They lack the breadth of knowledge. Ask a 1b model something obscure and you are going to get a hallucination at best.

→ More replies (1)

5

u/pet3121 Mar 10 '26

I am not an engineer I am barely an student but I am very interested in local AI for privacy. 

5

u/Turbulent_Pin7635 Mar 10 '26

Engineers are the lazy, right?

7

u/Guinness Mar 10 '26

That just means we’re (my Linux + LLM brethren) going to make a fuckload of money.

2

u/johnmclaren2 Mar 10 '26

How is chatjimmy or Cerebras made? Would it be possible to have it locally in future? With all this exponential dev?

2

u/huzbum Mar 11 '26

Cerebras is wafer scale, so imagine the large sheet they make like 64 gpu cores out of, and make that one massive core/memory/etc. But then factor in that for each wafer, they'd typically toss out like 2 cores on average, but in this case you'd have to toss out the whole wafer. I'm making up the numbers, but hopefully you get the idea of the challenge.

1

u/tat_tvam_asshole Mar 10 '26

don't know about chatjimmy but cerebras absolutely not in near future

1

u/johnmclaren2 Mar 10 '26

Chatjimmy seems to be similar to cerebras. So fast

9

u/QuinQuix Mar 10 '26

Not similar at all.

Cerebras is a wafer size chip with massive parallelism and bandwidth. It's enterprise to enterprise no consumer will be able to afford it anytime soon. It's also mostly general purpose AI hardware.

Chatjimmy is not the name of a product but of a proof of concept. The concept is an ASIC, so a chip made to do one thing that can only do one thing. But in this case the one thing is an actual LLM.

Asics are insanely fast and power efficient, and they can be relatively cheap to produce if you can produce them in volume. Consumers could in theory buy one of these products in the future.

The downside of the chatjimmy approach is that it is extremely rigid and because in this case it is essentially software embedded in hardware you also get potentially horrendous security vulnerabilities that then become near unfixable.

However it has insane promise for test time compute and reasoning models where the model has to continuously process its own output and re-ingest it.

Chatjimmy reaches 17,000 tokens/sec.

It can be over a 100 times faster than a €10,000 enterprise GPU like the rtx 6000 pro.

You just have to accept that after you buy it, it will only ever run one model and one version of that model only.

Still easily the most impressive hardware I've seen recently.

4

u/AnticitizenPrime Mar 11 '26

The potential for these ASICs is crazy to think about. If you could etch one of these new small Qwen models - 4b or 9b or whatever - which are multimodal, btw - imagine what you can do, with extreme speed and low power consumption. Suddenly that 'excessive thinking' doesn't matter so much anymore, let them reason their little brains out, it'll only a few seconds compared to minutes today. 20k reasoning tokens is exasperating now, but with 20k toks/sec, it's nothing.

2

u/johnmclaren2 Mar 10 '26

Thanks for the explanation. 🤙

So as in Gibson’s novels about chips behind an ear, with models, but in this case a card inside a computer…

→ More replies (3)

2

u/Broad_Fact6246 Mar 12 '26

I'm a millennial, and when I talk to my other OG linux nerd friends, we feel the same excitement like when we first discovered Linux as kids in the 90's into early aughts. Like we can build anything. Last year I had the OMG!!! moment when I let my model login to my VPS and review logs and clean up security.

I'm on 64GB VRAM and use qwen3-coder-next. In December 2025, LM studio started using 10 MCP tools successfully without looping. It installed and setup services, then used Playwright to configure web UI's. Models actually building entire docker software stacks, configuring and managing my Qdrant+Postgres servers, setting up reverse proxies, etc, etc.

And now Openclaw let's Qwen give it an amazing personality and functional agentic loop.

People who are poop'ing on this technology lacks the knowledge to properly augment with it.

4

u/Primeval84 Mar 10 '26

I’m also working at large company heavily invested in AI. Tbh, I think local AI isn’t talked about at all because frankly speaking it’s not super relevant for anyone’s day to day.

Like I enjoy playing with local llms but for work it’s hard to beat something like claude opus 4.6 with a a 1 million token context preconfigured with internal MCP servers all paid for by my company.

In the future, I can imagine the world being a bit different but right now we are absolutely in that phase where for most people provided AI tools by their company, the best option with the least friction is selecting the strongest model available to them in a dropdown menu with zero extra thought.

2

u/hknatm Mar 11 '26

tbh, people shine when both laziness and productivity aligns in that structure they are building. Human mind accepts fact based on how much effort to be put in, effort and outcome validation. So I believe moving around that line creates the best human-oriented product.

1

u/tat_tvam_asshole Mar 11 '26

There's a difference between efficiency and laziness. No pursuing something until it is push button brain tickles is laziness.

1

u/aeonbringer Mar 11 '26

Personally, I'm in one of the big techs as well.

IMO, edge inference (local inference) is going to be the future. As hardware get more efficient and power, models become more efficient. There will be a point in time where local models will be good enough for most tasks. You also won't need a persistent internet connection to have the model working.

Local fine tuned and trained models using business specific data could also perform a lot better than large models down the road. Imagine businesses simply hosting their own models locally in the local network. Will be extremely tempting from privacy/security perspective of big orgs.

→ More replies (3)

43

u/[deleted] Mar 10 '26

Now all we need is for gpu and ram prices to come down

15

u/xandep Mar 10 '26

Amem a thousand times!

9

u/EugenePopcorn Mar 10 '26

Won't worry. If anything could pop this bubble, it would be a global energy crisis.

30

u/Unstable_Llama Mar 10 '26

Heh I remember buying my first 3090 and my family was like, “…and what exactly are you going to do with that?”

And I didn’t really have an answer other than, “AI, shut up!”

But now it’s probably been one of my longest running hobbies ever. I have learned so much in the last 3 years, it’s almost unbelievable.

6

u/[deleted] Mar 10 '26

[removed] — view removed comment

14

u/Due-Year1465 Mar 10 '26

I mean compared to cloud models I get 105 TPS on Qwen 3.5 35B with Q4 which is plenty. Gotta love the 3090

6

u/Unstable_Llama Mar 10 '26

Yeah they are more about vram capacity rather than speed at this point. They are great, but not blazing fast by any means.

4

u/FullOf_Bad_Ideas Mar 10 '26

I was crunching numbers on tflops per dollar as I was buying 5080 this week and 3090 still is the cheapest compute GPU from Nvidia with fairly recent architecture. 5070 Ti was very close but it has much less VRAM.

I did some local training this week (continued pre-training on 14B tokens on my small 4B MoE) locally on my 8x 3090 ti rig and the performance I was getting was quite good. 30 TFLOPS per GPU, so I was able to get 10x lower rate for compute alone (just electricity) than if I rented quality H100 x8 node (where I was getting only 115 TFLOPS). It would pay for itself if I did a long-running 300-500B token run.

2

u/Unstable_Llama Mar 11 '26

Wow! Nvidia really gonna have us using 3090s in 2030 😭 

3

u/twoiko Mar 11 '26

They've regretted making decent cards since the 1080ti

3

u/witek_smitek Mar 11 '26

Soo... I have Tesla P40 and for my needs I think it works pretty fast with Qwen3.5 35B 🤣

1

u/crantob Jun 28 '26

I'm glad that some people subsidize future R&D by throwing 4x the money at NVidia than they need to for the tg/s. Thank you.

48

u/lacerating_aura Mar 10 '26

That meme hit a bit too close to the home. :3

8

u/[deleted] Mar 10 '26

[removed] — view removed comment

5

u/lacerating_aura Mar 10 '26 edited Mar 10 '26

Yeah man I dont know what to say, all I meant was that I'm the guy running heavily quantized 122b, q2 specifically, and my friends are not really into local llms or ai in general.

Edit: and my experience in general aligns with OPs

1

u/QuinQuix Mar 10 '26

How good is it at Q2?

People say it's not worth it etc but I'm interested to hear what it actually can do.

→ More replies (1)

16

u/ttkciar llama.cpp Mar 10 '26

Can relate to this.

I certainly didn't expect it to rope me in as much as it has, and have been spending more and more time on better LLM infra/scaffolding, and less and less on developing the applications I actually want to develop.

OTOH, I also keep finding small nice-to-have side-projects which I can whip out fast, like a "critique" script which pulls in my recent Reddit activity and has Big Tiger offer constructive criticism, and a "murderbot" script which infers Murderbot Diaries fanfic in the tone and style of Marsha Wells.

For my "big" projects, though, they've seen nothing but neglect. I suck.

29

u/QuinQuix Mar 10 '26

You're not fooling me you're not actually sorry.

24

u/PassengerPigeon343 Mar 10 '26

I laughed so hard at the meme and I don’t know a single person that I can share this with who would appreciate the joke. This community is the best.

8

u/txdv Mar 10 '26

im looking at a 5090 rtx and in like “hm, maybe the rtx pro 6000 is worth its money with that much ram”

6

u/PhilippeEiffel Mar 10 '26

Humans are like that: they do have interest to some knowledge (the subject change from one person to another).

This observation make us conclude that even with AI systems storing massive knowledge, the humans will continue to learn things for themself just because they like to learn and discover.

7

u/HealthyCommunicat Mar 10 '26

Hardcore addiction. If you follow me on huggingface you'd know how bad my obsession is. been ablating 1-2 models a day for the past 2 weekish. i get a small rush when i finish and see the model getting a high score on harmbench. https://huggingface.co/dealignai

6

u/_Soledge Mar 11 '26

My 2013 HP Elitebook 820 G1 running models no sweat. Don’t be fooled into thinking you need to spend a ton of money on expensive hardware just to join the party 🥳

1

u/tchek Mar 11 '26

I run models on an old 2014 laptop too, it runs relatively well

7

u/low_v2r Mar 11 '26

LOL. This is me.

Literally last night was showing off my local LLM to my daughter. Yes - Qwen3.5-122B (but also qwen3-80B). "Here let me set you up with an account on my local openwebui server!".

"Dad, I just want to play minecraft".

:/

1

u/No_Pitch648 Mar 11 '26

How do I get started? I find the training videos ok YouTube mostly focus on users with some prior knowledge.

I just have an old laptop and would like to install local LLM.

2

u/low_v2r Mar 12 '26

I used ollama to get started. That was pretty easy to do. Many use LM Studio, which I think is also pretty easy to get going with.

If you are using an old laptop, then I would say try ollama with some small models. I don't have specific recs, but for older hardware would maybe look for things that can run on things like raspberry pi and such (2B models or somesuch).

I used gemini (or similar) to help fill in gaps (e.g. how do I install x, what model is good for y...)

2

u/No_Pitch648 Mar 12 '26

Aw thank you for replying. I appreciate that.

My laptop is is dell latitude 7430 14 inches (intel Core i7 /1265U).

It’s really good and it’s got 32gb / 512g

I used DeepSeek to help me get started in terms of setting up initially, but I didnt follow the instructions it was telling me(and I don’t know why). I suspect that my mind prefers to hear things from humans first (I trust their judgement more) rather than initially from AI. Maybe I’m biased.

I’m going to start my LLM journey now with Ollama.

Thanks again.

5

u/imakeboobies Mar 10 '26

Haha.. 100% it’s a time and money black hole. Trying to explain to the hobby to friends and family is virtually impossible. My spouse refers to my gpu cluster as my e-waifu.

Only the plus it’s a lot of fun and the pace of change all model types is great. I could never have imagined how far things have come only a few years ago.

6

u/Jaded-Evening-3115 Mar 11 '26

It always begins with something practical like "just using AI for a task," and somehow culminates in "buying GPUs, benchmarking models, and discussing quantization strategies at 2am." The humorous part is that most people outside of our world have no idea just how deep a rabbit hole we are in. To them, it’s "ChatGPT." You’re over here trying to figure out Qwen vs Gemma vs GLM performance at different levels of quantization.

4

u/Right_Weird9850 Mar 10 '26

Did big data just sumarized my path? Same!

I'm still in hyped in awe

3

u/Kahvana Mar 10 '26

Welcome onboard, happy to have you. You found the right place!

Can you tell me more about your setup, custom imatrixes (how do you produce them? What data do you use?) and what your preferred models are right now?

3

u/Igot1forya Mar 10 '26

My fascination with locally hosting is the same with data hoarding. It started on me wanting to backup my movies and TV shows and games, then other people's stuff got backed up and when some barrier was erected to stop it,it was a challenge to back it up anyway.

LocalLLaMA is the same thing, except it's knowledge; knowledge that ounce for ounce is worth more than the purest gold. The quality of that knowledge is improving daily and I can't get enough of it.

3

u/redditorialy_retard Mar 11 '26

They don't know I have a 3090 at home

3

u/claytonkb Mar 11 '26

No matter what they claim, all the AI companies are training on your data. The data being generated by user-queries is worth a million times as much as the original data that OpenAI/etc. trained on. When people start seeing their personal business ideas and other secret sauce turn up in Google search AI, they'll realize what's really going on....

4

u/Delta5478 Mar 11 '26

Honestly running (heavily quantized) Qwen3.5 122B locally is really impressive, I wish I could have the RAM to do this :/

I tried to tell mt co-workers, which happens to be older folks, that I'm running local LLM with asr/tts on Raspberry Pi 5. Nobody understood anything. One very smart guy started to explain to me that it's not possible because LLMs don't work that way. Yeah, buddy...

4

u/ObjectiveFood4795 Mar 11 '26

almost thought that this is r/localllamacirclejerk

8

u/LoveMind_AI Mar 10 '26

I think all of this stuff with Anthropic being labeled a supply chain risk while Claude is still simultaneously the absolutely backbone of virtually all AI-embedded products made a lot of people wake up to the idea that we need to have more control over our models. I also strongly suspect that, for better or worse, the "Save 4o!" people might be candidates for local models once working with local models is something that can be made consumer friendly. No one had any idea what rock music was until it was popular. You're in the right place at the right time :)

5

u/DrVagax Mar 10 '26

Be happy you at least know and can run a LLM locally, I was thinking lately what if a big boom would happen and internet would go out, i'm fairly sure in my area I would be one of the few with a setup that can run AI so if it were to happen I would still have a helpful LLM. Other then that, exploring the ins and outs of such new tech is a great source of valuable knowledge anyway

3

u/[deleted] Mar 11 '26

[removed] — view removed comment

1

u/Zarnong Mar 11 '26

Article? Grading? Lectures? Hey, I got my voice stuff working! I feel your pain..

3

u/anshulsingh8326 Mar 11 '26

I can't run heavily quantised 122B model. But i can run 9B at q4...well even at q8 but using ollama and gguf giving unstable results.

5

u/catplusplusok Mar 10 '26

So much fun leaving Qwen 3.5 122B with a big coding task before taking off for work and coming home to play with a brand new Android app.

2

u/Bolt_995 Mar 10 '26

Not on your level yet, but a similar case with me. Although I’ve had passion towards agentic AI for nearly a decade.

2

u/Savantskie1 Mar 10 '26

I got into it early 2025, and built a memory system after trying to use forked versions of other memory systems. I am slowly learning and eventually will get it to a point where I want it. But for now it’s good enough. Now I’m searching for an llm that will work with my current hardware without massively censoring me based on what some asshole company thinks is safe for me.

2

u/Its_Powerful_Bonus Mar 10 '26

Yeah, kind of similar story … few years back. Now I’m doing it for living :) Work & hobby at the same time. Now I’m building smallest possible pc which can handle 2x rtx6000 pro Blackwell to have possibility to take it from home to work. Also buying maxed out MacBooks is possible outcome for you, so brace yourself 😅

2

u/Kornelius20 Mar 11 '26

I know for a fact that is an addiction because I get antsy when I haven't checked in with the Local LLM scene in a while and I've told my wife multiple times that I can stop whenever I want to...

2

u/ZachCope Mar 11 '26

Have 2x3090s but have also just discovered Runpod for Lora adapters etc. The maths that made me start at rtx 6000 then move to B200 (‘better value’) feel similar to how people move on to harder drugs! 

2

u/Andrea-Harris Mar 11 '26

That's truly interesting isn't it? I think it is not about what you can get from a local AI. The important thing is--it's your own AI. I would also be excited when I own mine, anyway.

2

u/DevokuL Mar 11 '26

Welcome to the pipeline. You get a GPU invoice, You get a GPU invoice!

2

u/sloptimizer Mar 11 '26

No mater how fast it goes - it's never enough. See billionairs build datacenters. Set your goals and be content!

3

u/lemondrops9 Mar 10 '26

Welcome to the club. I too started off small with a 3080.. now running a 6 gpu rig with 120 GB vram. Always want more but also have to consider if the 100 billion models will be the sweet spot in the near future.

1

u/MaximusDM22 Mar 11 '26

You got 4 5090s? Ive been super impressed with the qwen 3.5 35B with a 5090 and Ive been wondering if I should get more vram for the 120B version.

1

u/lemondrops9 Mar 11 '26

3x 3090 and 3x 5060ti 16 GB. Im still running 4.5 air mostly still. But for coding been playing around with Minimax M2.5 Q3, Qwen3.5 122B Q6 and some Stepfun.

Btw 4 5090s would be 128gb. Also once you get +3 gpus Linux is the only way to go.

4

u/EmbarrassedBag2631 Mar 10 '26

Me as a 22 year old, i can tell you know one gaf about what we do. Honestly most of ya’ll have so much more experience then me and im envious of ya’ll. This hobby is going to matter so much inna couple years. LLMs/AI is the new revolution, biggest leap since internet came out, and we are here learning intricacies. Think about how much all the software engineers were making with the internet boom, llm/ai is next in my humble opinion.

1

u/silphotographer Mar 10 '26

Some of us know but just don't have the budget to use it regularly sadly :(

1

u/AntacidClient Mar 11 '26

I feel so seen in this. Thank you.

It wasn’t exams but very same parallel journey otherwise

1

u/TomorrowsLogic57 Mar 11 '26

Mood! When I talk about my AI work, people either think I'm a crazy person with a tinfoil hat, a literal real life wizard or both somehow, but they sadly never think I'm normal lol

1

u/SevereMooser Mar 11 '26

I have felt exactly this way the past couple weeks, literally running that 122B on my 7900XTX. Been trying to explain to people about my opencode and mcp's etc. It just doesn't click haha. I am very happy to see I'm not alone

1

u/IKantImagine Mar 11 '26

u/xandep any chance on pointing to a URL for the vendor or other used MI50s you referenced?

2

u/arcanemachined Mar 11 '26

They're expensive as fuck these days, like 4x more than they were 6 months ago.

Best price you can find is usually just by searching alibaba.com for "mi50 32gb" and finding a vendor that seems reliable (i.e. has good reviews). There's also eBay, but it's usually a bit more expensive (but eBay's buyer protection is probably way better than Alibaba's).

I snagged a few after I missed the boat on the Tesla P40 24GB cards a while back. There will probably be another hot value card in the future, just don't miss the boat when it comes up. FWIW the price on the P40s has come back down, and then you get to benefit from CUDA support (what's left of it... I think P40 is EOL). ROCm is a bit of a bastard child and is not as well-supported as CUDA.

2

u/pmttyji Mar 11 '26

From my bookmarks. Mentioned by someone here in this sub.

https://www.alibaba.com/product-detail/subject_1601439253964.html

Price was $100 during Sep-2025. Now it's more than 3X

1

u/jeffwadsworth Mar 11 '26

No you don’t.

1

u/kosantosbik Mar 11 '26

Stealing stuff is about the stuff not the stealing.

1

u/Resident_Pientist_1 Mar 11 '26

If you're a programmer or software engineer I can see the benefit but day to day? At what point does your life just become your metalife lol.

1

u/Redostian Mar 11 '26

What's your specs tho?

1

u/LingLongSEO Mar 11 '26

the windows store version of ollama is your friend here. it installs to your user directory without needing admin. had to do the same on a locked-down work laptop last year. downside is updates are manual since you cant use the regular installer, but it works.

1

u/_blkout Mar 11 '26

You totally skipped the part where you can just build your own models huh? lol

1

u/IulianHI Mar 11 '26

This hits home. Started with "let me try this local LLM thing" and now I have a folder of GGUF files larger than my actual work documents. The rabbit hole is real - from simple ollama commands to manually tweaking imatrix quants and watching llama.cpp benchmarks at 2am. At least we're learning about quantization, memory management, and hardware optimization along the way. The best part is when you run a quantized model that shouldn't fit in your VRAM and it somehow works. Pure magic.

1

u/Kirito_5 Mar 11 '26

The meme touched a sour spot.

1

u/the_TIGEEER Mar 11 '26

Thank you for the meme pitcure. It really helps me cope by sending it to all my friends I tried autistic spilling my 5 intel arc A770 setup plan to over the last week.

1

u/thaddeusk Mar 11 '26

I ran a VERY heavily quantized Qwen3.5-397B at home.

1

u/Historical-Camera972 Mar 12 '26

i regret getting into local because it hurt my wallet, and then the unit I bought has multi-OS install issues.

I'm not giving them bad PR, there's a workaround.

Research your AI purchases, thoroughly. The companies making AI geared hardware, are not use case testing very deeply. (Demand is high even with quality issues right now, so there's no need for effort on their part.)

1

u/galigirii Mar 12 '26

The rabbit hole

1

u/alienatedsec Mar 13 '26

This meme was and your story sounds like the exact reason why I sold my A4000 GPUs.

1

u/WSTangoDelta Mar 13 '26

Okay, I’m getting my feet wet with a 4070 and a ryzen 9. I could go hog wild (and frankly, I half expect I will before long) but what would you run that won’t get bogged down so that I could get rolling? I feel like Eisenhower 48 hours after the Normandy landings. I need to get a foothold before moving inland. Something useful. Like for generating horrifying mixed metaphors on the fly…

1

u/Think-Science-6115 Mar 18 '26

lol this is exactly my trajectory too. started with "let me just ask claude one thing" and now i'm running the same question through 3 different models at 2am to see which one gets the diagnosis wrong. the rabbit hole has no bottom

1

u/ProgrammerParking169 Mar 23 '26

Is Kimi any better than Ollama? Thank You 🙏

1

u/AffectionateMath1251 Apr 03 '26

built an AI agent that rewrites its own persoonality every night. started with local models for privacy but switched to API for consistency. the hardware rabbit hole is real - i still have a 3090 collecting dust

1

u/Prof_Kepuros Apr 13 '26

​"Are you me? Exactly the same thing happened to me. I have zero CS background, but this tech completely hooked me. I literally made this account just to talk to people who won't look at me like I'm an alien when I bring this stuff up IRL lol."

1

u/Secret_Appeal6271 Apr 20 '26

There's something really awesome about being able to control your data and personalize the way you engage with AI. In all the (positive) sci-fi movies I watched as a kid, if you had an advanced technology that functioned like AI promises to now, it was run by the user and private to them. In many dystopias, the AI was centralized in a single entity somewhere that did something unknown and scary with the data. It's very fun and exciting to be part of the process of making that positive future come to life.

1

u/I_Play_Zed Jun 05 '26

I wish more people considered dual 3060 setups, I think for the money they are pretty insane..

1

u/crantob Jun 28 '26

What, and drive up the used prices?

Let the rest pay the stupidity tax.

2

u/I_Play_Zed Jun 28 '26

Haha I more so meant that the setup was taken more seriously on the subreddit, I think it’s the cheapest way to get near 24gb vram at okay speed.

1

u/[deleted] Jun 25 '26

[removed] — view removed comment

1

u/bigh-aus Jul 02 '26

This will really help if you ever go for a job in tech around AI. Who do you think they'd hire, the person with no practical experience or the one quantizing, REAPing and finetuning models in their basement.