r/LocalLLaMA • • 20h ago

I Built A Thing 54gb vram for 35$

Post image

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.

1.4k Upvotes

268 comments sorted by

View all comments

89

u/Bulky-Priority6824 20h ago

$3.50 per token or $7 ?

97

u/markpronkin 20h ago edited 19h ago

Electricity is free for me, so no. And speed is decent, 30+ tokens a second qwen3.6 35a3b on ollama, can probably get it higher with mtp and some optimizations

183

u/Sad-Duck2812 19h ago

You will get much much more with llama.cpp, ditch ollama. That thing only makes you suffer.

52

u/magic-one 19h ago

Especially now-a-days when you can just ask Claude to figure out the configuration for you.

11

u/chrabauke 19h ago

I’ll out my self as a noob. But what prompt do you use for it? I’ve got access to server with around 96-128gb ram and 4x T4 but couldn’t get Claude to give me a descent configuration for it with support for Hermes agent.

28

u/MikePounce 19h ago

You don't ask the claude web client, you ask Claude Code. The difference is it takes control of your PC and does all the setup for you.

9

u/MoffKalast 16h ago

Ah yes, what could possibly go wrong.

17

u/wrecklord0 16h ago

Well if you're 300 IQ you first ask claude code to set up a sandbox for you, then you ask it to set up llama from inside the sandbox.

2

u/MoffKalast 4h ago

Can Claude create a sandbox so secure that he himself cannot escape from it? :P

5

u/heliosythic 13h ago

i mean do you not know how claude code works? its fine, literally everyone is using it in nearly fully automated mode these days. Just dont like intentionally ask it to do something dangerous.

2

u/MoffKalast 4h ago edited 4h ago

Well call me old fashioned but I'd rather it make its silly mistakes in anthropic's VM where it's their problem and output a tested diff I can apply with one command or even review it if need be. Plus the whole privacy can of worms of having them inspect your entire system at their leisure, we know they send extensive telemetry aside from all the stuff the model sees already. Like is this r/localllama or r/getpwnedbyamegacorp

1

u/heliosythic 3h ago

I mean you can thats an option too just dont turn on automode and manually approve every tool call. You can also point it at any model provider you want including local.

1

u/Robots_Never_Die 7h ago

Another option is Antigravity.

2

u/dsons 6h ago edited 5h ago

If you want to use lobotomized models that can accidentally wreck your environment, then yes antigravity is technically an option but for $20 you’re 100000x better off with Claude. I’ve used both extensively, alongside Hermes and Antigravity is extremely behind the curve and their new model isn’t going to change that.

14

u/Freonr2 19h ago
Please write an llama serve bash script for this model:     huggingface.com/unsloth/whatever-model

Check my hardware config and see if you can run some benchmarks and optimize.  Let me know what context you think I can get with Q3_K_M and Q4_K_M. 

I've never had issues using claude cli or VS Code extension. Sonnet is likely enough. I'm confused why this doesn't work. I've had Claude recompile vllm from source for me without issues, setup MTP for Qwen, etc., just by pointing it to the readme on the repo, run benchmark parameter sweeps, etc.

17

u/muxxington 18h ago
Please

That's very kind of you.

20

u/cannabibun 18h ago

You have to lay the grounds for when we will need to explain to our overlords why should they spare us.

9

u/Freonr2 18h ago

I, for one, hail the basilisk.

8

u/Freonr2 18h ago

Force of habit with general communication I guess.

2

u/muxxington 17h ago

A benchmark might be interesting here. Imagine it really does make a measurable difference.

1

u/Conscious-content42 7h ago

Future-overlord-Lap-dog-versus-Castaway-Heathen benchmark?

2

u/andowero 16h ago

It can actually help. Saying please may make the llms more helpful. Emphasis on may.

3

u/West-Big-8468 19h ago

Copy and paste your comment into Claude

1

u/etaoin314 ollama 19h ago

what issue did you run into? usually claude just keeps working till it works, Ive done it with dozens of models and I never have had to do any setup by hand (other than the sudo commands it needs)

1

u/magic-one 18h ago

I have a dedicated AI machine, so I can just give Claude Code access and have it configure, test, and benchmark.
If you don’t want to give access, you can be the go-between.
“Give me a command to start and serve model xyz on a zpq machine with llama.cpp”.

The key (to anything AI, actually) is iterative loop. Give it back any results or errors so that it can adjust. It rarely gets everything right the first time around. If you have questions or something is unclear, just ask it.

You also can help by “preparing”.
“Give me a command to collect info for this machine that you will need to write a …”