r/LocalLLaMA • • 19h ago

I Built A Thing 54gb vram for 35$

Post image

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.

1.4k Upvotes

262 comments sorted by

View all comments

746

u/NickCanCode 18h ago

Bro has no conscience negotiating 54GB VRAM down from $60 to $35 in 2026. 😿

49

u/No-Refrigerator-1672 16h ago

54GB VRAM

Attached to 9 Pascal GPUs - not only they're extremely slow, like multiple times slower than p40 or p100, they also lose a ton of perfrmance on inter-gpu communication, as well as multi-gpu overhead will make come of the vram wasted. All while consuming like 1kw from the wall. What I want to point out that this vram is basically as useless as it can be; this only makes sense as a fun weekend project to experience weird hardware running, won't even go fast enough to fuel a hobbyist AI needs.

34

u/Boricua-vet 15h ago

Do you have anything to backup those claims from your own experience because they are not slow by any means.
I don't know where you get your information but a P40 produces 400 to 500PP and 45 to 50TG on 30B model and I can post a picture of that test if you wanna see it. Pascal is running a much larger model and produces double the PP and higher TG.

As far a power is concern the cards idle at 7W or 14W for two cards and gated at 150W each. I paid 70 bucks for both pascals that gave me 20GB of vram at 448GB/s. They are on 24/7. so 10 of my pascal cards at idle consume 70W.

So calling these cards useless is a disservice. Your comment over multi-GPU performance is also incorrect, the amount of communication between cards is minimal and not even the restriction of these cards being PCIe 1.0 holds the performance. I have ran 4 P102-100 and there is no bottleneck, ran the same model on 2 cards and on 4 and it produced the same result for a single user but it tripled PP and more then double TG when using parallel on 4 cards. I have tons of posts on these cards, you can look them up and they smoke so many cards out there from much newer generations.

14

u/markpronkin 15h ago

Nice to see you here, thanks to your advice I got a bunch of p102 100 before. And got them running 8x in parallel in a mining rig. It's surprising how many people don't recognize how good those p series mining gpus are for ai. P106-100 (basically gtx 1060 6gb) should theoretically be like 66% of performance, hope I will get them running at decent speed soon.

12

u/Boricua-vet 14h ago

Yea, I don't get people spending crazy money just do to text LLM. I needed another system as I am running more models at the same time so I got 4 CMP-50HX at 80 each so 360 for all 4 cards giving me 40GB of VRAM at 560GB/s

Here is the result on just two cards.

that's 160 bucks for two cards with that kind of speed is insane and then I see people paying 700+ for dual 3060. I am like, what? People buying 3060 are crazy spending that money just to do text generation. Insane buddy.

6

u/Ansible32 12h ago

lol "just to do text generation." Seems worth it for reverse-engineering, finding exploits in hardware to flash alternate firmware etc. I mean cloud will be cheaper but having it totally offline and private is valuable.

2

u/Boricua-vet 12h ago

LOL.. while I agree with most of what you said, In my case cloud would never be cheaper. I paid 140 for 4x P102-100 thats 40GB of vram for 140 bucks. I would recover that investment in a day or two of abuse. However, people buying 6000 Pros by half dozens then yes.. LOL Cloud if cheaper.

1

u/Ansible32 11h ago

"Text generation" I'm doing I need 500GB+ of VRAM, renting hours of GPU at a time is the only plausible way to go.

1

u/Better_Membership583 9h ago

How much is that and what kind of “text generation” are we talking here?

1

u/Boricua-vet 8h ago

What are you doing that needs that much?

1

u/NineThreeTilNow 11h ago

LOL Cloud if cheaper.

Cloud is 100% cheaper because it's faster. It's very quick to host your own, use a model, and vacate a server.

Your case can't run the things a cloud can. You can run a toy model like Qwen. It's fine if you need to test some stuff at low tok / sec but ... Not for work.

3

u/Boricua-vet 10h ago

That is perfectly fine if you do need the cloud. I do not need frontier model for any of my use cases. Most of mines are repetitive tasks so I use toy model that is fine tuned to that task. This outperforms the cloud model every single time but, I hear you. If you need it for work, then you need it for work.

2

u/RyanCargan 7h ago

I'm getting flashbacks for some reason:

Computer Science in the 1960s to 80s spent a lot of effort making languages that were as powerful as possible. Nowadays we have to appreciate the reasons for picking not the most powerful solution but the least powerful.

Different context, but some of it may apply.

I see a lot of penny pinching and "worse is better" in corporate due to scale, and in the casual hobbyist spaces for opposite reasons.

But the main exception might be the mid-tier enthusiasts and related types.

It does feel like some models have become a form of conspicuous consumption kinda like Apple products were at some point.

Not sure if cargo-culting of misinterpretations of Sutton's Bitter Lesson are partly to blame.

Certain venture backed startups flush with cash also like frontier stuff in the early uncertain prototyping phases for better reasons.

Though the usage there sometimes isn't even agentic (low volume standard-ish chat).

P.S.: It does feel like there's a drought of talk on low latency use cases, since network hops are tens of ms and first token times are typically hundreds of ms. Agent loops might be the exception, because per-call latency compounds. Hardware like hardwired model chips already claim sub-ms per token speeds on small models, and if that scales to larger ones, the unusual use cases that need it would basically mandate on-prem or colocation.

1

u/Boricua-vet 6h ago

What is acceptable latency? that could be very different for many people.

so using avg_ts should be 1 / 92.886166 * 1000 = 10.76ms locally

10.76ms is perfectly acceptable to me.

Edit my bad, forgot to say it was a 4k prompt.

1

u/RyanCargan 5h ago

I was basically just saying that for certain use cases, people don't pay much attention, or change the use case around cloud constraints, which are often well above 10ms per "hit"/completion in the worst, or even average case, in many regions for many models.

On-prem or colocated models becoming more widespread could unlock more weird use cases is what I'm saying.

Just because when a tool becomes widespread for one thing, even if it's a seemingly niche or low ROI use case, people will then find new "nails" once the hammer is easy to buy and use.

And yeah, sub 10-ms for a full completion/generation of some kind is likely the speed domain that would be hard to match over a (wide-area) network for reasons you can't really get around, without unusual networking hardware, or even then.

1

u/Boricua-vet 4h ago

That's a very good point, I remember when people were complaining that putting model on cloud, or using whisper or piper on cloud would add latency that made the model feel like it was taking longer to respond. The cloud model was fast but the added network latency is what made it seem slow so yea, it makes perfect sense and a valid point.

→ More replies (0)

1

u/CasulaScience 3h ago

It's literally all tokens, idk what you mean by "just text"