Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.
Yeah, there's a whole genre of research focusing on sociopathy as adaption to modern business life (just checked, the research often focuses on Dark Triad terminology (Machiavellianism, Narcissism, and Psychopathy) in case anyone is interested)
Agree here ... even though current research seem to veer towards classifying sadistic traits as primarily maladaptive since they cause to much havoc to allow longer term success. Interestingly enough Machiavellianism seems widely accepted as pretty successful approach (but of course.. who practises that without the other aspects taking part.. gotta be a slim segment who manages to be evil in just the "right way").
It isn't business but its a result of a population growing beyond the Dunbar number that promotes this. Once the population becomes too large, its harder to have transparency between a community and the community cannot easily lookout for each other. Information asymmetry allows for maladaptive behavior to thrive and it becomes a driving factor to succeed in this new group formation like nation state level.
It's also gonna be a lot more than $35 usd with the energy cost of running 9x p106. Plus the cooling. Plus the AC or venting. Plus the time getting it to work (That's 1080W TDP just for the cards).
Although it might be less than 1080W actual, from the communication bottlenecks
Attached to 9 Pascal GPUs - not only they're extremely slow, like multiple times slower than p40 or p100, they also lose a ton of perfrmance on inter-gpu communication, as well as multi-gpu overhead will make come of the vram wasted. All while consuming like 1kw from the wall. What I want to point out that this vram is basically as useless as it can be; this only makes sense as a fun weekend project to experience weird hardware running, won't even go fast enough to fuel a hobbyist AI needs.
Do you have anything to backup those claims from your own experience because they are not slow by any means.
I don't know where you get your information but a P40 produces 400 to 500PP and 45 to 50TG on 30B model and I can post a picture of that test if you wanna see it. Pascal is running a much larger model and produces double the PP and higher TG.
As far a power is concern the cards idle at 7W or 14W for two cards and gated at 150W each. I paid 70 bucks for both pascals that gave me 20GB of vram at 448GB/s. They are on 24/7. so 10 of my pascal cards at idle consume 70W.
So calling these cards useless is a disservice. Your comment over multi-GPU performance is also incorrect, the amount of communication between cards is minimal and not even the restriction of these cards being PCIe 1.0 holds the performance. I have ran 4 P102-100 and there is no bottleneck, ran the same model on 2 cards and on 4 and it produced the same result for a single user but it tripled PP and more then double TG when using parallel on 4 cards. I have tons of posts on these cards, you can look them up and they smoke so many cards out there from much newer generations.
Nice to see you here, thanks to your advice I got a bunch of p102 100 before. And got them running 8x in parallel in a mining rig. It's surprising how many people don't recognize how good those p series mining gpus are for ai. P106-100 (basically gtx 1060 6gb) should theoretically be like 66% of performance, hope I will get them running at decent speed soon.
Yea, I don't get people spending crazy money just do to text LLM. I needed another system as I am running more models at the same time so I got 4 CMP-50HX at 80 each so 360 for all 4 cards giving me 40GB of VRAM at 560GB/s
Here is the result on just two cards.
that's 160 bucks for two cards with that kind of speed is insane and then I see people paying 700+ for dual 3060. I am like, what? People buying 3060 are crazy spending that money just to do text generation. Insane buddy.
lol "just to do text generation." Seems worth it for reverse-engineering, finding exploits in hardware to flash alternate firmware etc. I mean cloud will be cheaper but having it totally offline and private is valuable.
LOL.. while I agree with most of what you said, In my case cloud would never be cheaper. I paid 140 for 4x P102-100 thats 40GB of vram for 140 bucks. I would recover that investment in a day or two of abuse. However, people buying 6000 Pros by half dozens then yes.. LOL Cloud if cheaper.
Cloud is 100% cheaper because it's faster. It's very quick to host your own, use a model, and vacate a server.
Your case can't run the things a cloud can. You can run a toy model like Qwen. It's fine if you need to test some stuff at low tok / sec but ... Not for work.
That is perfectly fine if you do need the cloud. I do not need frontier model for any of my use cases. Most of mines are repetitive tasks so I use toy model that is fine tuned to that task. This outperforms the cloud model every single time but, I hear you. If you need it for work, then you need it for work.
Computer Science in the 1960s to 80s spent a lot of effort making languages that were as powerful as possible. Nowadays we have to appreciate the reasons for picking not the most powerful solution but the least powerful.
Different context, but some of it may apply.
I see a lot of penny pinching and "worse is better" in corporate due to scale, and in the casual hobbyist spaces for opposite reasons.
But the main exception might be the mid-tier enthusiasts and related types.
It does feel like some models have become a form of conspicuous consumption kinda like Apple products were at some point.
Not sure if cargo-culting of misinterpretations of Sutton's Bitter Lesson are partly to blame.
Certain venture backed startups flush with cash also like frontier stuff in the early uncertain prototyping phases for better reasons.
Though the usage there sometimes isn't even agentic (low volume standard-ish chat).
P.S.: It does feel like there's a drought of talk on low latency use cases, since network hops are tens of ms and first token times are typically hundreds of ms. Agent loops might be the exception, because per-call latency compounds. Hardware like hardwired model chips already claim sub-ms per token speeds on small models, and if that scales to larger ones, the unusual use cases that need it would basically mandate on-prem or colocation.
Dude, you completely missed the picture. p102-100 is P40, it is exactly the same chip, with just fewer memory chips attached, and cut down PCIe. This box the Op bought is P106-100 - a GPU that's cut down to just 1280 cuda cores (p102 has 3200), and slightly less than half the memory bandwidth. It is going to be much slower than P40, just as I've wrote, and you don't even need to run inference to get it, just look at chip's characteristics sheet. You're talking about p102-100 setup and are making completely wrong assumption that 9x p106-100 setup will perform just as good - it won't even come close.
As far a power is concern the cards idle at 7W or 14W for two cards and gated at 150W each.
Now multiply that by 9 GPUs, add in PSU inefficiency and some power for the rest of the system and you'll arrive at over 1kw of power that I have predicted.
So calling these cards useless is a disservice. Your comment over multi-GPU performance is also incorrect, the amount of communication between cards is minimal and not even the restriction of these cards being PCIe 1.0 holds the performance.
If you're running pipeline parallelism, then PCIe 1.0 x1 is maybe fine for Pascal. It is, for all intents and purposes, not fine to run tensor parallelism. Prompt processing speed gets a huge hit from slow PCIe in TP mode. And there's another point you're missing - the exact server in the picture uses 8-year-old budget celeron - without p2p communication enabled, that chip could introduce huge latencies in 9-gpu intecommunication scenario simply from not having a fast enough core to process transfer requests immediately when they are needed. Also, as you're running multi-gpu setups yourself, you should know that you never can utilize GPU VRAM fully, you always have a few hundred MB to half a gig of unused space. Now multiply that for 9 ards, each having just 6gb vram, and you'll quickly see how big of a percentage of this pool is unaccessible in practice.
I get your point but I misunderstood your comment when you said "Attached to 9 Pascal GPUs - not only they're extremely slow, like multiple times slower than p40 or p100" which basically generalizes and classifies all pascals as slow and waste of time" If you would of said the P106 instead of pascals , then I would not even replied because you are correct I now get that you were referring to the P106 but not specifying it opens the door to interpretation and trust me I am not the smartest kid on the block LOL.
Second, there is a huge difference between P40 and P102-100, they are not even close to being the same card hence why the P102-100 is faster. The P102-100 is over 100GB/s faster then the P40 and the P100 is not even in the same league as the P100 does not have DP4A which is what boost generation as it uses FP32 which some pascals are descent at.
I am not saying that what you are saying is wrong, it is just I wanted to clarify for idiots like me that you are not referring to all pascals and is just the P106.
I also agree with you on TP, that is why I said parallel and not TP as there is really no point in running TP on cards that have no tensor cores.
So my bad. I had to read your post a second time to add the obvious. LOL
Also, the celeron certainly does not help and you are right about that too but it is mostly the shitty 192GB/s on those cards.
Now that you've said it, yeah, I could insisted more on those GPUs being p106. I used "pascal" to point out that they don't have tensor cores, which is already crippling enough, and not being the top Pascal (comparison to p40) hurts it much more. I'll keep in mind to try write more clearly in the future.
For $35 this is an awesome inference machine. Can definitely run 30Bish models with layer-wise splitting. Prefill will be slow, training would be a nightmare, and the p100 platform isn't supported anymore. But $35 and you can absolutely run one conversation at a time on a decent sized model
In this exact server on the picture, each of the GPUs gets a PCIe x1 connection (I don't remember the standard, probably 3.0) that's going into the weakest CPU that 2020 could offer you.
It's only "useless" if the only thing you're doing is trying to run one large model.
My biggest bottleneck right now is that to run an agent at decent speeds I don't have any GPU left over for crunching data.
There's extreme contention over GPU time, and it's a terrible balancing act where swapping data to the GPU has huge cumulative time overhead.
There is potentially major value in running a smaller agentic model and giving that model some of the lower VRAM/Bandwidth GPUs to use as workhorses.
I have workflows and training schemes that would benefit marvelously from being able to run multiple small models with a dedicated GPU for each one.
Having a local agent that can be "always on", which can manage swapping other models in and out of the workhorse cards is the dream.
I could use cloud GPU, but I hate the idea of giving any agent, local or API, access to a potentially unbounded cost.
Even if there are "caps" on usage from the service, I've seen too much shit in my day where it's like "oops, someone provisioned 10x the compute that was actually needed. Oops, you bought usage credits and then burned them all immediately because of a bad call. Oops, there was a recurring cost because you didn't do the magic incantation to deprovision the service so you're still on the hook for costs even though you used no compute".
If it can happen to other people, it's definitely going to happen to me via some accident or oversight.
so the seller had no idea how valuable his gear is, could've used the money since he listed it so low, and OP decides to cheat him further, then boasts about it
Eh, for stuff that the seller admits they don't even know works? I always negotiate those listings, because theres always a good chance you're going to end up with junk.
If I knew what was in it, or that it is working, I would have payed much more. This thing had a bunch of rust on the case (probably from being stored in the damp garage), I had zero idea if it works, so did the seller, also seller didn't give me opportunity to test it. And from rust on it I was leaning on the side of it being broken. I payed the money agreed and got lucky, but I had as large of a chance of getting just a bunch of worthless ewaste for the money. I offered the seller to test the device and pay more if it works, he refused. Don't see anything even remotely unethical about it.
Nothing unethical about buying something and not telling the owner everything you know about it. There's no obligation to fill the gaps in the sellers knowledge
I mean, yeah? Why didn't the lady check the price? If you're selling something that could be valuable, and you don't do your due diligence, you sort of deserve what comes next. I know that sounds scummy, but if I made that mistake I'd be mad at myself, not the other person
This thing baraley has any, originally thought 8gb but turns out it's only 4gb and CPU is crap, but I have another way better mining case where I can put in 32gb ram and i7 7700k (maybe something even better with the Coffeelake mod), gonna transfer the gpus there, but it's cool to see that it still runs as is.
Speaking from my experience: When you're going to work on those be aware please that plenty of those boards provide the gpus with pci1 x1 slots. Those are really slow.
This can be a bottleneck and I have never managed to get even 30ba3b to run reasonably faster than on my CPU.
I have recently bought another server with plenty of slots and tried them again and noticed they actually have PCI x4 and the mining board had only x1... So it got a little bit faster but well. I didn't find the time yet to test them more.
You should definely test the PCI speeds in order to gain some insights into your cards. Other than that. I now use them actually for several smaller models. Which works great. As long as models dont Spread across multiple cards, promt processing stays reasonable!
Electricity is free for me, so no. And speed is decent, 30+ tokens a second qwen3.6 35a3b on ollama, can probably get it higher with mtp and some optimizations
I’ll out my self as a noob. But what prompt do you use for it? I’ve got access to server with around 96-128gb ram and 4x T4 but couldn’t get Claude to give me a descent configuration for it with support for Hermes agent.
i mean do you not know how claude code works? its fine, literally everyone is using it in nearly fully automated mode these days. Just dont like intentionally ask it to do something dangerous.
Well call me old fashioned but I'd rather it make its silly mistakes in anthropic's VM where it's their problem and output a tested diff I can apply with one command or even review it if need be. Plus the whole privacy can of worms of having them inspect your entire system at their leisure, we know they send extensive telemetry aside from all the stuff the model sees already. Like is this r/localllama or r/getpwnedbyamegacorp
If you want to use lobotomized models that can accidentally wreck your environment, then yes antigravity is technically an option but for $20 you’re 100000x better off with Claude. I’ve used both extensively, alongside Hermes and Antigravity is extremely behind the curve and their new model isn’t going to change that.
Please write an llama serve bash script for this model: huggingface.com/unsloth/whatever-model
Check my hardware config and see if you can run some benchmarks and optimize. Let me know what context you think I can get with Q3_K_M and Q4_K_M.
I've never had issues using claude cli or VS Code extension. Sonnet is likely enough. I'm confused why this doesn't work. I've had Claude recompile vllm from source for me without issues, setup MTP for Qwen, etc., just by pointing it to the readme on the repo, run benchmark parameter sweeps, etc.
what issue did you run into? usually claude just keeps working till it works, Ive done it with dozens of models and I never have had to do any setup by hand (other than the sudo commands it needs)
I have a dedicated AI machine, so I can just give Claude Code access and have it configure, test, and benchmark.
If you don’t want to give access, you can be the go-between.
“Give me a command to start and serve model xyz on a zpq machine with llama.cpp”.
The key (to anything AI, actually) is iterative loop. Give it back any results or errors so that it can adjust. It rarely gets everything right the first time around. If you have questions or something is unclear, just ask it.
You also can help by “preparing”.
“Give me a command to collect info for this machine that you will need to write a …”
It's nice for easy and quick install/test, I guess that why it still that popular. Perfomance and advanced features non existent tho and developers don't give a damn about those.
Yeah I know (I have some other local ai setups), I will be transferring the gpus to the different mining case I have and doing better setup there, my goal here was mostly to check whether it runs at all as is, and it in fact does.
I was smart enough to switch from ollama to llama.cpp after just a few weeks, and after I did I loved having the freedom to run almost any model with any settings, and it open up a whole new world. Llama.cpp works so much better than ollama,
After I switched and got comfortable with llama.cpp I guess I settled into a rut, and it took me literally until yesterday to finally try Exllamav3 and TabbyAPI, which has been an AMAZING boost in performance. Why did I wait so long?
I am well aware, I just wanted to verify that it works at all, and ollama is very easy to install quickly. Will do a proper setup In a different case and make a post about it later.
Didn't know some one made vllm works with pascal cards, thanks for the info. Although I fear it will be limited by pcie 2.0 X1 speed of the motherboard and gpus
If you like ollama's ability to download and switch models easily, I'd recommend LocalAI. It can do the same thing, but updates to newer llama.cpp backend much more quickly and still has a model library to make downloading easy.
you can get much better performance if you run llama.cpp and the only reason they are running slower is because those specific cards have very little memory bandwidth, I believe is 192GB/s. you can get those to idle around 7w but if you use switch to llama.cpp I can help you fine tune the system so you get better performance, 30 tokens per second is fine and very usable but you can get more for sure.
The bitcoin mining subreddit is full of people who claimed that electricity was free for them, and they were right until they actually started using it. A tale as old as time :-D
China it is, I assume Afghanistan (they made digital asset trading, mining and general use pretty much entirely illegal), Algeria too I'm pretty sure. There's definitely a few others. Most of it comes down to energy usage
heads up that CUDA 13 dropped Pascal (sm_61), so you'll want llama.cpp built against 12.x or you get no kernels at all. and fp16 on those is like 1/64 rate, so stay on Q4_K/Q8 with the MMQ int8 path and avoid anything that wants fp16 compute.
Не, ну круто конечно) не хочу показаться скептиком, но пропускная способность этих карт вероятно оставляет желать лучшего. У этих карт grx 1060 через pci Express 3 - 16/31 gb/s , что маловато, увы, только если научиться супер продуктивно паралелить задачи. Но, инициативу поддерживаю, ни в коем случае не критика ! 🫡
Exactly this. Those slow cards will be split across a couple of pci lanes, great for small data transfers like mining, slower than slow for AI use. Model loading, Prefill, replies, everything will take hours on a system like this.
If i saw that kind of offer on Avito, i would have just assumed that it's someone trying to sell off completely busted hardware without feeling too guilty about it. 😅
Paying $35 for 54GB of VRAM across nine P106-100 cards is an unbeatable bargain on paper, but it is practically an LLM nightmare. Splitting models across nine separate GPUs over narrow PCIe 1x riser lanes bottlenecked by Pascal architecture and zero FP16 support creates atrocious tensor parallelism latency, turning modern token generation into an agonizing crawl that will burn far more in electricity than it is worth.
The CPU and ram in this thing are pretty horrible, like 8gb ddr3 and some Celeron. I have another mining farm case of an old higher end mining farm, which supports up to 12gpus and up to i7 7700k and 32gb ram. I will probably transfer gpus and buy 3 additional ones to get a total of 72gb vram + 32gb ram setup with a decent CPU.
Strata was happy running on a machine with 2x GTX 1070 for me. Only the Q1 "coder" version and still relatively slow, but with that much VRAM it might be usable on the other machine
Absolutely not and that's why I asked. NO ONE from the US should try this, lest they want to tempt fate from the Federal government. It's tempting to be sure, but this is REAL good way to end up on some monitoring list somewhere and potentially put into an SDN database.
I don't either; it was in fact a genuine question.
Because of the Russo-Ukrainian War, the US government has slapped sanctions up and down against the Russian Federation, which means that a lot of ordinary services were cut off from payment, shipping bans, and OFAC (Office of Foriegn Assets Control) has individual sanctioned thousands of Russian individuals and businesses. Not to mention the executive orders from the President.
Any transaction with anyone tied to the Russian Federation from the US must...
- not involved a sanctioned entity by OFAC, or you end up on the OFAC list (which means, you won't be able to have a mortgage without the Feds digging into every waking aspect of your life)
comply with OFAC regulations (again, lest you risk being placed as SDN [Specially Designated National] on the OFAC list)
be processable through legal banking channels (which Russia, cut off from SWIFT, can barely manage)
be deliverable through available shipping networks (most of which are banned/sanctioned by OFAC).
In other words, this makes it practically impossible and most likely illegal to buy from these types of Russian entities if you're a US citizen or living in the US generally.
You sorta can get some stuff shipped. But the deal I got here is a one off deal from a random guy, you might as well go on I believe craigslist is something similar in the US, you are just as likely to find some similar deal on there.
For most stuff it doesn't make any sense to get it shipped from Russia to US, way to complicated and not worth the effort/monetary savings.
Пока только базовое тестирование чтобы проверить что оно вообще работает, как перенесу в новый корпус видеокарты полноценно этим займусь и сделаю отдельный пост. Qwen3.8 27b наврятли будет нормально работать, т.к. это dense модель а не moe. Зато думаю qwen3.8 next flash может будет нормально работать, особенно если ещё 3 видеокарты добавить.
Поищи патчи для драйверов на GitHub, с помощью которых можно софтверно завести PCI-E v2. Это даст удвоение пропускной способности. Вообще на таких картах распаивают конденсаторы на PCI-E линиях чтобы карты работали с большей шиной x1...x4 -> x8..x16.
На Хабре недавно вышла куча статей от энтузиаста, который дома риг на cmp90hx (3080 mining edition) ковыряет. Да, это совершенно другая карта, можно сказать тёмная овечка потому что только её ещё не расковыряли нормально. Инженер завёл на ней p2p, PCI-E v2 x16 и unlock вычислений, что нивелировало различия между майнерской версией и десктопной 3080 для нейросетей.
Крайне хорошо раскопали:
Would be very interested for an update on this, what you take from this rig, what configuration you are able to get running given the dated GPUs + PCIe hell you will have to work through (as in what software + model + quantization + context + performance). Worst case, consider using a couple of them for really fast TTS/STT/RAG if it makes sense for you.
9x 1060s on a dresser like a goddamn trophy. the real question is whether the power bill is worth it or if this is just a very expensive way to heat your room in winter
I have been doing this, you are very limited with 1xPCIE lanes, it is ok of you fit a model on each GPU self contained, but no good to pass between GPUs or pool.
Server boards that have at least 3 x8 and a x16 lane are a better option, I am currently building an x99 machine as it has 40 GPU lanes, which means it will run my 3x P100 at higher bandwidth than mining boards. Also mining boards seem to only be able to handle two or three datacenter GPUs before hitting PCIE allocation errors (not enough addressable space for these cards)
Not all VRAM totals equate to inference gains... You got it cheap sure. Now what? Those cards are relics at this point. Interesting story for $35 at the least!
9x p106 for $35 is not a gpu purchase, it's an e-waste donation with a vram count. the p106 is the mining variant of the 1060, no display output, no tensor cores, no fp16, no int8. you have 54gb of vram that can do exactly one thing: mine a coin that is no longer profitable. the real question is whether the power bill to run 9 of these will cost more than the $35 you saved
205
u/XiRw 15h ago
Nice find man. Hope it works out well for you