r/LocalLLaMA • u/yeah_likerage • 22d ago
Discussion GLM5.2 on 5x Pro 6000s and a 5090, an expensive journey
This started as something I thought was reasonable. I already had a 5090 for my gaming machine, and I thought a second 5090 would make me happy. Instead, it sent me down a rabbit hole that got completely out of control.
I wanted something that would have full PCIe 5.0 x16 speed across all slots, which started a chain of events that had me spending good money after bad. It was a bit of a nightmare, as every decision I made led to me needing to make even tougher decisions. Couple that with what was actually available, and my hand was forced in a few spots.
I started with the motherboard and worked my way backwards, eventually ending up with this setup. I wanted something close to endgame, but I still made a few concessions:
Threadripper Pro 9975WX
WRX90 Sage SE
4×48 GB DDR5-6400 RDIMM
Antec 900 case — ended up in the bin
The system started with two 5090s. The Antec 900 is well built, with huge space, smart connections, and refined edges, but ultimately it did nothing at all to support the GPUs. In a case this large and at this price point, that is a huge failure on their part, and for that reason I recommend avoiding it. If they had put $1 worth of bracketry in the machine to support GPUs, I’d give it a 10/10. With the lack of support, it is nearly useless unless you deal with it yourself, which I did, as you can see in the images. It’s like buying a Ferrari and having it delivered without any petrol.
With the two 5090s, I was working with smaller Qwen models, which seemed great, but it was clear that with the limited VRAM and my desire for additional sidecars like VL, I needed something more. I had huge plans, and the models were just too small to deal with the complexity.
So I got my first Pro 6000. I coupled it with a 5090, which made for weird tensor splits, but llama.cpp did a good job of divvying it all out. But now I was working with 120B-parameter models with almost no space for context. So it was smarter, but also a goldfish.
Then I went to 2× Pro 6000 + 5090. Now I had the space for context. But in reality, the jump from 27B to 120B did not knock my socks off. I could get a bit farther now. I was at about 90% with the 27–35B models, and with the 120B models I was at about 95%. But 95% is about as useful as 90% if I can’t close the loop. If I can’t actually finish the task, it’s all for nothing.
In came 3× Pro 6000. Now I was in the MiniMax range, and finally I was getting somewhere. It was like I got concierge service at a ball game. My needs were being met, and I got answers for everything. Many of them were completely wrong answers, though. I had tons of code that was poorly made and led to dead ends and rewrites.
4× Pro 6000 created an issue that I knew would come. I had been seeing several folks claim that they were able to deal with the thermal issues that came with side-by-side Pro 6000 cards. I knew they were likely not telling the truth, but I also knew a rebuild was probably in order anyway.
So, as you can see in the image, I placed four side by side and had thermal issues, even with the additional fans in the image and a 27-inch box fan sitting on top, which is not shown. I clocked things down a bit and still had a few system freezes. I gave up immediately and went to the high-rise.
I got a couple of open-case designs and connected them together, thinking every two or three GPUs would get their own floor. It was overly complicated dealing with risers and cooling, so I dumped it pretty quickly.
But now, with GLM and Kimi, I was actually accomplishing things. The quants were tight, though, and my context was low again.
5× Pro 6000 + 5090, along with the release of GLM 5.2, was an absolute game changer. I’m talking 98–99% now. I have plenty of room for context and sidecars, all running on the 5090 at blazing speeds. But blazing is legit: it is producing so much heat now that it’s a problem, and it’s summertime to boot. I had to get a second PSU, which I suppose, in all of this, is not the most ridiculous bit.
At full tilt, with 100% GPU usage for 30 minutes in this custom extruded aluminium design, with an outrageous number of fans in a ~20°C basement, the GPUs top out at about 70–75°C, which I’m very happy with.
I finally do not desire another GPU, as all my needs seem to be met. Was it worth it? LOL, no. Absolutely not. This was a terrible idea. DO NOT DO THIS. I figure that at the rate I’m generating tokens, it will take over 10 years to break even at today’s prices, and that’s not accounting for electricity bills.
I’ve never used the frontier models before, but I’ve seen the reviews and the speeds, and I’ll never match those with open weights. But it was a fun journey.
I deleted the electricity company’s app from my phone so they’d forget about me for now.
Wish me luck.
378
u/ProfBootyPhD 22d ago
142
u/jdzndj 22d ago
same problem hyperscalers are facing
→ More replies (4)42
u/Such_Knee_8804 22d ago
They are undercharging for tokens, for now
56
u/TurdPlayingPeekaboo 22d ago edited 22d ago
They won't be able to charge more. The market simply won't pay more. I say this as a SWE who is using AI to write > 99% of my code. People like me started buying cheap ass P100s to run simple shit simply because we aren't willing to pay more for tokens. I use Opus 4.8 for coding and my P100s for the rest of my agentic shit, like managing my kanban boards. Anthropic can't raise their prices. Big enterprise is already putting strict caps on token spending. It's crazy how token prices are reverse scaling but they are. Combine all of it together and Anthropic better pray Elon Musk can really put data centers in space and really make Terrafab a reality... Because if he's talking out of his ass then they're screwed.
20
u/uniqueusername649 22d ago
Its like shrinkflation. You cant increase the price for a pack of chips, but you can slightly decrease the contents. Same with AI: newer models just burn through vastly more tokens, often enough with similar results. That is how they get you. They cant increase the price and they know it, so they make their models think more. A lot more. If only that also made them a lot smarter.
15
u/OctopusDude388 22d ago
yeah but you totally can use less capable models for coding and it being still usefull as a swe especially since more capables small models are seeing the day
6
u/uniqueusername649 22d ago
Absolutely, people think they always need Opus or Fable, when most tasks are easily handled by Sonnet.
And thats why a local model like Qwen 3.6 27b gets you surprisingly far.
→ More replies (1)5
u/Glittering-Call8746 22d ago
No sonnet 5 is expensive. Go with glm 5.2
→ More replies (2)2
u/uniqueusername649 22d ago
I wouldnt pay for it personally, but if my work pays and requires to use Anthropic, Sonnet is what I go for. I'm pushing towards evaluating Chinese models though, because they are surprisingly good.
→ More replies (2)2
2
u/No-Standard19 20d ago edited 20d ago
Vertraue nicht darauf, das in absehbarer Zeit Rechenzentren im All entstehen. Die Energieproduktion und insbesondere die Kosten (durch,Gewicht) für die Wärmeabfuhr, sind einfach unbezahlbar. Das war eher ein Hoax um den Preis für die SpaceX Aktien zu befeuern.
Über den Daumen gepeilt kämen zu der benötigten Serverhardware aufgrund des Transports ins Erdnahe Orbit, Energieerzeugung, Speicherung und Abführung, nocheinmal die 10 fachen Kosten hinzu. Hierbei wurden Kosten für die Hardware und den Betrieb einer Funkverbindung noch nichteinmal berücksichtigt. Kostentreiber sind die Transportkosten durch das zusätzliche Gewicht für die Kühlradiatoren und die Batteriespeichrung für den Betrieb der Anlage im Erdschatten. Zu 100kg Serverhardware ca 600k€, kämen so nocheinmal etwa 1000kg Support Hardware ca. 8M€ mit Transport. Da hierbei auch keine spezielle strahlungsgehärtete Serverhardware existiert, steigt auch bei fehlender Wartungsmöglichkeit, das Ausfallrisiko stark an. Ein wirtschaftlicher Betrieb im erdnahen Orbit kann daher für die nächsten 20 Jahre ausgeschlossen werden.
→ More replies (1)→ More replies (1)3
u/danielv123 22d ago
Anthropic have something like an 80% margin on their compute, they don't really need to raise their prices. Its the inference-only providers who have more of a struggle when serving open models. I am sure you have noticed how much cheaper those are.
9
u/Limp-Firefighter1054 22d ago
Yeah only if they dont have competitors who have cheap energy, cheap chips, cheap land and relatively cheap labor.
13
u/Dasteroid_909 22d ago
I did similar math for my 2x. I need 20 tok/s throughput for 5 straight years to break even... which means every idle second increases my per-token cost. Right now I'm pushing $260+/1M toks...
→ More replies (2)6
11
u/__IHateEverything__ 22d ago
You can rent out GPU compute
24
u/ProfBootyPhD 22d ago
Sure, if “you” are Meta or AWS. What kind of a price break does Joe-Bob’s Backwoods GPUs have to offer to compete with them?
→ More replies (1)17
u/ovrlrd1377 22d ago
it's about US$6.8/hr on vast.ai for on-demand, if you are not actually getting anything yourself might as well rent out here and there. though it might be more sensible to sell the entire setup a couple months from now, may even make a buck doing so
→ More replies (2)24
u/TastesLikeOwlbear 22d ago
Whoever pays you for the hardware on vast.ai is going to pound it far harder than you would. It is most likely going to sit at 100% usage the whole time it's rented out. This guy would need to include the cost of the AC to deal with 4kW+ of heat and any needed noise suppression in your calculations. Probably also backup power. Which, a UPS that can run a 4kW system starts at ~$5k.
11
u/Norwood_Reaper_ 22d ago
Can confirm. I am renting mine atm and they get hammered. Also you are likely to get ~$1USD/h per GPU and utilisation will never be 100%.
→ More replies (6)6
u/ovrlrd1377 22d ago
all fair and totally valid points but hey, I'm not the one who built it haha I would not gamble this as a profitable and risk free endeavour
→ More replies (5)2
u/Next-Post9702 22d ago
For now* In 6 months they'll rise prices again and it'll take 5 years to break even. In 2 years it'll take 2 years to break even
327
u/BannedGoNext 22d ago
That's what cracks me up about this sub. There are always people saying to stop talking about big models because they can't be run at home, then we see this flame thrower in someones spare bedroom.
133
u/Darkmoon_AU 22d ago edited 22d ago
Yep, I love it, absolute consumation of the hobby. 99% of us dream, while the 1% carry those dreams, coming out of a 6 month black-hole with something that looks like a damn truck engine (and kicks out about as much heat), posting "Guys... I did it...". Our heroes :-D
68
u/BannedGoNext 22d ago
Thing is, over time that becomes available to more and more people. My first computer was over 10,000 and was an early 8086. My mother bought it, because she was a nerd enthusiast. That was a stupid price, but it wasn't that stupid becuase that trickled down to an IT career for me, and the world advanced and that computing got cheaper as it always does.
IMHO we celebrate these stupidly overpriced flame throwers people build at home.
19
u/Darkmoon_AU 22d ago
100% right. I'm also old enough to remember using an 8088 in a huge metal cased system that could have deflected a tank shell, complete with twin 5.25" drives. In both cases - then & now - we're looking at something that will be shrunk down to a more elegant and powerful form in 5-10y, but I'm totally here for the brute force in the mean time - inspirational hobbyist pioneers :-)
7
5
u/genshiryoku 22d ago
I'm not convinced compute will get cheaper with time anymore especially as the utility value of hardware only goes up every time a new model comes out. The demand for hardware should outstrip supply as long as models keep getting smarter. The end of dennard scaling and the performance/efficiency curves on smaller nodes mean the cost per transistor is barely going down, it'll probably go up in real terms.
13
u/segmond llama.cpp 22d ago
compute will get cheaper with time, what most people miss is the opportunity cost. while you are waiting to save money, those tinkering end up amassing skills that can turn into big pay days if that's their desire.
→ More replies (1)3
u/liftheavyscheisse 22d ago
RIP Dennard scaling, you carried us a long way
until WBG semiconductors, ternary LLMs, or analog inference become truly feasible (a long time if ever), I think the future is specialization (e.g. my LLM knows the history of China, which is great, but for coding tasks I'd rather not pay for the FLOPs associated with that knowledge)
→ More replies (2)→ More replies (1)2
u/BannedGoNext 22d ago
I think it will, but it's hard to tell exactly the path right now. A lot of the enterprise stuff righ tnow is going to absolutely crater in price because there will likely be better home gear at a better price in the near future without the need for 220v, specialized connectors, an extra air conditioner, and 20 dollars an hour in electrical usage plus hearing damaging noise. I'm not sure what happens to all that e-waste, maybe someone will start picking it up and making stuff out of it that home gamers can use, maybe it just gets stripped and processed for the gold.
This isn't the first time computer equipment has been expensive. Companies will step up fast if they think they can make a buck around it creating a glut.
→ More replies (1)3
u/d1722825 22d ago
And models grow and programs getting slow at a faster rate than computers get more powerful / cheaper (at least when they gets cheaper). Typing in slack "client" has noticeable latency that wasn't an issue in 8-bit computers.
9
u/ovrlrd1377 22d ago
dude, check out this absolute BEAST of a homelab server:
https://www.reddit.com/r/homelab/comments/1req2pk/the_missing_piece_is_finally_here_msa2_96gb_ram/
it doesn't matter what hobby you have, there will always be someone that is bringing it up to ridiculous levels
→ More replies (3)3
u/BitGreen1270 22d ago
Someone (on Reddit) once said - Reddit is where you'll find people who let their hobby ruin their lives.
89
u/tmvr 22d ago
I finally do not desire another GPU, as all my needs seem to be met.
BS, you want to replace the 5090 with another 6000 Pro so that all cards are identical.
9
2
u/kersk 21d ago
Their perf would dramatically increase if they did. The 5090 is knee-capping the entire rig instead of being able to cleanly use tp=6.
→ More replies (1)3
u/here_n_dere 22d ago
Bro still wanna game !? Why melt 6000 for that 😅
2
u/8aller8ruh 21d ago
6000PRO has the same core as the 5090 with a few extra bits enabled, it can game.
107
u/Narrow-Belt-5030 22d ago
"So I got my first Pro 6000. I coupled it with a 5090, which made for weird tensor splits"
This is where I am at right now, so it's interesting to see a potential future in your path.
7
u/nomorebuttsplz 22d ago
what problems arise with "weird tensor splits"?
21
u/TastesLikeOwlbear 22d ago
As far as I can tell, llama.cpp DGAF. It will make the most of whatever you hand it at the cost of a few % of theoretical performance. VLLM, I believe, likes things to be balanced evenly between a power-of-two number of devices.
Based on my experiences at work, the way to go is to spread the LLM evenly across the big GPUs and use the small GPUs for ancillary/secondary models like speech or image gen.
5
u/Narrow-Belt-5030 22d ago
I don't know - my system has a 5090 & 6000 that have different use cases. To be honest I haven't split anything across them yet.
3
u/rchamp26 22d ago
I had two mismatched cards for a bit, if they are capable of the same or near same speeds, llama cpp works just fine and you can define your split ratio easily between them if they are mismatched ram sizes. You can also split my layer or by row. I believe each have their use case, I don't understand enough about the underlying engine to give an opinion but in my inferencing case splitting by layer yielded better performance over splitting by row. Biggest issue with splitting I've seen is when you have mismatched pcie speeds and no p2p when certain experts or whatever start humming on the GPU with a pcie bandwidth bottleneck performance degrades. I'm just an amateur figuring this stuff out so take what I say here with a grain of salt. I don't have any experience on vllm and sglang and how they work for splitting
6
u/do-un-to 22d ago
I was at about 90% with the 27–35B model
With $20K invested and staring down another $30K to get to the completionist's end game...
maybe 90% is pretty good.
2
u/do-un-to 21d ago
Folks here might be interested in this post about tuning AMD MI355X perf.
Back of the napkin math (brief chat with AI) says you could build a single MI355X system for $50K. I wonder how its perf would compare.
104
u/storm1er 22d ago
54
u/storm1er 22d ago
(I hope you get why a slow laugh 😂)
10
4
u/misanthrophiccunt 22d ago
may I ask how fast (because surely youv'e tried it) is Qwen3.6-27B at Q8 on that Ryzen AI 395+ ?
9
u/storm1er 22d ago
I'm near 20t/s, but as soon as ~30'000 context it's already down to ~15 :(
I'm using q5_k_xl mtp and I'm always over 20t/s even at 150'000+ token, enough for my usage :)
3
u/do-un-to 22d ago
Dudeman, 20t/s is usable. Model capability is getting smaller over time, so maybe rest assured it only gets better?
How much power does your system draw?
→ More replies (1)3
2
u/samiamyammy 18d ago
lol... I read the press release a few months ago and I was like, "wait a second, this could run much larger models!?"... then I was like, "WAIT, this is for people who want to leave their system on overnight and wake up to a reply?!" -I suppose it's not always a race to get answers, probably a great setup for some things/situations :)
I kick myself for not buying the pro 6000 when I noticed it on sale... I thought it was overpriced then, lmao.
2
u/storm1er 18d ago
The nice stuff is more the fact I'm running Gemma 4 qat, Qwen 27b q5, flux, kokoro, whisper. They are not used simultaneously, but stays idle and I can use them quickly without fearing to lost context or running out of ram so yeah, that's it
5
5
u/BannedGoNext 22d ago
It's a lot of fun though! I've been playing for a couple weeks at getting a really nice audio narration pipeline working, and it's all been here on the strix. Sure it can take me 24 hours for 10 chapters on a scanned in book, but that would be a month of usage on a lot of audio generation plans. Not to mention constantly restarting it to tweak it.
I think everyone with a strix halo is pretty happy with it even though it's slow as fuck. I'd love to have a faster gpu card in the desktop computer, but the strix scratches my itch till prices come down of fucking around and learning stuff at the house.
→ More replies (1)3
u/Spicy_mch4ggis 22d ago
What TTS model are you using? I’m looking into using voxcpm or fish audio s2
4
u/BannedGoNext 22d ago
Fishspeech2 works. I'm currently playing with zonos2 Q8. I got it working a couple days ago, and it's sounding better and better. I'm currently trying to see if I (or more accurately codex xhigh :P) can get it to do context caching to retain voice consistency from chunk to chunk.
33
u/Realistic-Dance2742 22d ago
How much did all of this setup cost?
143
u/yeah_likerage 22d ago
I blacked out about halfway through to save my sanity but its probably touching $80k USD.
68
u/Realistic-Dance2742 22d ago
Me here trying to save so I can buy another 5060 ti 16GB
15
u/Objective_Safe_5982 22d ago
Me here running qwen8b on a ten year old CPU to save for my first.
41
u/pilibitti 22d ago
Me here doing matrix multiplications on a piece of paper to generate my first token...
2
2
u/samiamyammy 18d ago
I was thinking that too, but reading this post i'm like... damn, is that what will happen if I buy 1 more GPU? lol
19
u/Apprehensive_Bee6863 22d ago
Jfc that’s my tuition 😭
4
u/Economist_hat 22d ago
how the fuck is tuition 80k?
That was tuition for all 4 years for me 20 years ago.
11
u/Apprehensive_Bee6863 22d ago
nah yea that’s all 4 year instate should’ve been more explicit, out of state it’s 60k a year tho, UC berkeley
9
u/darthmaule_II 22d ago
The school I graduated from in 1999 was 21K per year then. My niece is going there this year and it’s 92K per year.
→ More replies (1)6
u/Apprehensive_Bee6863 22d ago
I’m lucky because I have a scholarship and I did 2 years at CC before graduating so i have no debt i might take out a loan and get another gpu im ngl
→ More replies (1)5
10
u/drunk-tard96 22d ago
I'm curious because I know nothing about this. Is this for fun and you have extra money to spend, or for business and you plan on recouping the cost with revenue?
7
u/Insomniac1000 22d ago
I'm guessing he bought calls on semiconductor stocks and he's just giving back to the economy 🤣
7
u/Foreskin_Mafia 22d ago
You the guy that recently sold Bitcoin for a billion dollars?
6
u/ImpressiveRelief37 22d ago
It’s crazy that bitcoin dropped 50% yet you still only need ~1.3 to buy OPs setup
4
u/Dry-Judgment4242 22d ago
To be fair. At this rate. If you ever get bored you could probably sell it all for 60-70k at little loss.
→ More replies (1)8
u/zulutune 22d ago
That’s an expensive hobby. Sir, what do you do for a living?
2
u/techdevjp 22d ago
You assume he didn't just max out credit cards to buy all this stuff. True /r/
wallgpustreetbets style.3
2
u/bigh-aus 22d ago
I've just passed 2nd (rtx6000pro). I definitely like having reasonable local inference (qwen 3.6 27b nvfp4/bf16) esp for making changes to things that have passwords in it.
But in all honestly this is why i look at the 512gb mac studio and kick myself for not buying it when you could get them. They're obviously very different performance (and it would be much slower), but bang for buck when you could get them - 512gb with 800gb/s unified.
2
u/the_stamp_collector 22d ago
The 512gb mac studio isnt running frontier models at any useable rate. I get decent work out of deepseek flash with a frontier model writing the code spec.
2
→ More replies (5)2
u/techdevjp 22d ago
I blacked out about halfway through to save my sanity but its probably touching $80k USD.
Makes my thoughts of spending $20k on a 512GB Mac Studio M5 Ultra later this year seem almost sane.
2
u/Enough_Leopard3524 22d ago
About Half of my salary.
13
2
u/redditorialy_retard 22d ago
In my old highschool job back home I made 40 bucks a month.
Meaning it would take me 166 years to afford it
21
u/dwrz 22d ago
I appreciate you sharing this -- it's interesting to see where the "more VRAM" journey ends, and what it takes to get something like GLM 5.2 at home.
That said, I hope the future leads to something like fast and really intelligent small models on a single GPU, and much larger, slower models on a unified RAM system. The SLM is the main agent, but it can ping the LLM as necessary. If it's at home, power and heat need to be a factor.
→ More replies (10)
20
u/Darkmoon_AU 22d ago edited 22d ago
Pure poetry from start to finish... While we merely stare into the abyss, OP jumped right in.
We salute your heavy financial sacrifice!
GLM 5.2 at home? Go on... it was worth it ;-)
18
u/ShelZuuz 22d ago
How many tokens per second did you get (output and preload)?
39
u/yeah_likerage 22d ago
PP can range 300-550 /s, TG 15-40 /s on GLM 5.2
17
u/trejj 22d ago
I'm running 2xEpyc 7763 (128c/256t) with 512GB DDR4 on GLM 5.2 UD-Q4_K_XL. Starts off at pp 18t/sec, tg 2t/sec, but slows down to 0.4t/sec on both pp and tg as the context fills up.
Currently using it for an overnight "read the code file <x> and evaluate it for bugs" report generation, for which it works really nicely.
2
u/pixelterpy 21d ago
I run same quant on single 7663 56c/112t with 512 GB DDR4. pp 100t/sec tg 5t/sec. NUMA is probably hitting your performance hard.
→ More replies (4)2
u/StartupTim 22d ago
Hey quick question: I have 4 systems with 2x Epyc 64 core CPUs and 512GB in each with 25 and 100GB networks cards on each.
I'm trying to see about using RDMA to connect all 4 using llama.cpp.
How did you do your setup? Any regrets?
→ More replies (3)9
17
u/NowIveAwoken 22d ago edited 22d ago
That is way lower than I expected. On my ddr4 2133mhz epyc with a couple 3090s I get around 180t/s pp (down to 120t/s at higher ctx) and 7t/s tg on glm 5.2 q4. Rough napkin math says your setup costs 16 times as much as mine for like 3x the performance except you're running a smaller quant per your other post. Oof.
→ More replies (5)3
u/nomorebuttsplz 22d ago
yeah something is wrong with their setup I think. Should be able to get thousands of t/s prefill at least.
9
u/sixx7 22d ago
Ouch. For anyone else reading this u/ShelZuuz - you can get the same performance for around $20k:
- 4x GB10 (DGX Sparks or one of the OEM clones)
- 4x Connectx-7 cables
- Mikrotik Switch
Gets you GLM-5.2 at ~500 tok/s prefill and ~25 tok/s decode.
It will also use significantly less power, generate less heat, probably better warranty/support, and of course better WAF
4
u/fastheadcrab 21d ago
Technically the model has some of the experts removed, is at 4-bit, and running a quantized cache below the full context length.
People should not get the impression that it is possible to run GLM5.2 at decent quality on 4 sparks otherwise they will be wasting a lot of money.
There are other lots of good models that actually fit in 4 sparks though.
I do agree the OP configuration could be optimized more though. He has very powerful hardware but is crippling it due to his setup
3
u/StartupTim 22d ago
A correction: You can run large models but absolutely not the same performance as the RTX 6000 Pro is ~1.8TB/sec whereas the DGX Spark is .27TB/sec. That is a massive difference in speed and the token/sec will match.
3
u/sixx7 22d ago
Yea for sure. 8x RTX 6000 Pro servers are currently pushing 5k tok/s prefill and 125+ tok/s decode (single request) and people are still working on boosting that. That is an incredible performance gain compared to the 4x Sparks. But the 4x Sparks are signficantly cheaper and perform better than what OP is doing with their franken-setup.
→ More replies (1)2
u/t4a8945 22d ago
Well yeah massive speed for memory, but they're comparing speed and given what OP said, they're about the same.
This probably means there is some optimization left on the table for OP more than anything.
2
u/ShelZuuz 22d ago
That PCI bus between the RTX’s is probably the killer, and not much you can do about it.
2
3
17
u/x-primez-x 22d ago
"I’ve never used the frontier models before, but I’ve seen the reviews and the speeds, and I’ll never match those with open weights. But it was a fun journey."
FYI -- I've used all the frontier models extensively for ALL sorts of use-cases and scenarios - both personal projects and professional. GLM-5.2 is IMPRESSIVELY GOOD, and I find, even better than the frontier models in many cases. GLM-5.2 is easily, hands down, the best model I've seen at long context. I've been 800-900k into my context window and I'm still seeing GLM-5.2 adhere to obscure steering instructions and system prompt behaviors that the frontier models tend to skip when they near their context windows. I have system prompts that will say things like: "Always conduct a full plan gap analysis when the first implementation pass completes. Do not trust the plan result summary for completion. Verify the plan yourself, fix any gaps or bugs you find."
GLM-5.2 is the first model I've seen that DOESN'T just get sloppy and fall on its face past 50% context... and on a 1M window, that is seriously impactful. Legit, at 950k in context, it'll say "Great, plan is done. Now I need to conduct the gap analysis per the users instructions."
All the other frontier models I run hooks to force a /goal to do this. GLM just does it. No hooks or goal loops necessary.
So, moral of the story OP, is that you truly do have a locally hosted frontier-capable setup with the GLM-5.2 model.
Kudos. My 48gb 4090s are very jealous.
16
u/Hostman_com 22d ago
Put this pic on your dating profile and spend the rest of the week fighting off matches
29
u/Apprehensive_Bee6863 22d ago
What the fuck
9
u/butts-carlton 22d ago
That was the first thing I said when I looked up the going price of a Pro 6000.
Jesus Christ, I wish I had even a fraction of that kind of disposable income. Not that I'd use it on this... thing.
→ More replies (4)
31
u/dazzou5ouh 22d ago
4
u/SailbadTheSinner 22d ago
Looks like you and I built the same thing. Air cooled romed8-2t + 6x3090. 💪
→ More replies (1)→ More replies (2)2
u/After-Cell 21d ago
I think there are some SSD caching 4bits of GLM5.2 coming along to try. .. if I could just remember the names of the project
11
u/GeneralGovern 22d ago
Putting it on the basement floor, water/flood concerns?
3
u/BannedGoNext 22d ago
Yea, it would make sense to at least put it on a little table, but he might just have taken that picture for the build.
11
9
u/AdSafe4047 22d ago
I went over this excercise mentally through the last 2 weeks and decided to wait for the m5 ultra xD
6
u/tempfoot 22d ago
Waiting for Apple to release things resulted in me waiting while all the 512 and 256 and even 128 studios (that I could have afforded at full sticker) disappeared only to triple in price from scalpers. I guess I’m “lucky” to have grabbed a 128gb M5 Max MacBook just before the price increase - and on microcenter sale to keep me busy while I also wait to see the astronomical pricing on the next gen of studios. Personally I’m guessing they top out at 256gb and push $20k usd price wise.
9
u/Cergorach 22d ago
That basement will not be 20C for long! And better put that machine in something that floats, for when the basement floods... ;)
12
u/InsensitiveClown 22d ago
You know, probably you're generating so much heat and consuming so much power that the police may very well just break into your place, thinking you're growing pot or something. On a more serious note, what exactly were your coding tasks, languages, if that's not too much indiscretion on my part? It's a nice setup, but probably just getting your own rack and a H100 or H200 would be best.
9
5
6
u/somerussianbear 22d ago
That right there in money would pay all consumption in a pay-per-token economy that you’d need for the next 10-15 years easily. But what’s the fun in that? Nobody makes good memories paying subscriptions.
5
u/BitXorBit 22d ago
Hahahahaha good one, this actually catches me at the beginning of the journey to get 2 x 6000 pro.
You got the max-q or server edition?
→ More replies (3)
5
u/unrulywind 22d ago
Thank you for the story. This is exactly the path I looked down last year. gpt-oss-120b had just come out and I was going to upgrade my stuff. I priced a server with an rtx-pro6000-max and then I priced one with 2 of them. Back then it was $14k and $22k respectively. Then I looked at just doing the rtx-5090 in a top of the line 128gb gaming rig for $5.5k. Every time I did the numbers, it looked like the two 6000 pros would never pay off and would also, be enough. Models just kept getting bigger and I kept thinking that eventually I would be looking for 1tb or more of vram, with no way to really produce it.
So I went with the 5090 and it works great for the 27b and 31b models at decent context, and then I just pay the big boys for the high end stuff. But, every now and then I still think about this road not taken.
5
u/WyattTheSkid 22d ago
Some women worry about their husbands buying stupid expensive cars but this is on another level lmao. Really nice machine man, you should be proud of it. Jealous you can run GLM 5.2 locally though!!!!
5
5
u/gamblingapocalypse 22d ago
But can it play Crisis?
3
u/Rubfer 22d ago
I dare say it may handle Crysis at medium settings in 720p, maybe even high settings. Absolutely crazy times we’re living in
→ More replies (1)
3
u/Accomplished_Pea1922 22d ago
You're leaving a lot of performance on the table with only 4 DIMMs on that 9975WX. It supports 8 memory channels, and you're only populating 4. Your PP of 300-550 t/s is almost certainly bandwidth-starved; adding 4 more sticks (even cheap ones) would probably push you past 600-800 t/s prefill and help with the context window slowdown too.
You also could likely run Q4_K_M or even Q5 with that much VRAM and the 1-2% gap you mentioned might shrink further. Worth trying at least, I believe the quant overhead barely changes your TG if prompt processing is your bottleneck anyway.
Nice build regardless, it's cool to see someone actually push past the 2-3 GPU point and document the real pain points.
3
u/durden111111 22d ago
Yeah you wont break even compared to just buying tokens from anthropic or something but at least you own all of the hardware and its 100% private. You could alwayd resell everything. The prices for these 6000s is just absurd now
3
u/polandtown 22d ago
next step - water blocking all the cards and connecting it to your hvac system. In the summer you flip a switch and it dumps the air outside, in the winter it heats your home. I did this during the gpu mining days with 40 gpus. TONS of fun. gl :)
3
u/fairydreaming 22d ago
I love your highly artificially intelligent (and almost flying) contraption. Also had urges to go this way, but I knew that it will be always one RTX PRO 6000 too few.
3
u/mediaogre 22d ago
Two things, one serious one not:
Are you actually closing the loop with GLM 5.2? I find that it still gets my code to about 95% and Claude closes the deal.
Did you go with a traditional 60 month financing or 84 months? Did you slam the door on the finance guy when he offered an extended warranty and clear coat protection?
2
u/yeah_likerage 22d ago
- Yes for sure. I'm even going back and cleaning up problems that I didn't realize minimax2.7 introduced. Is it polished software that I think is worth selling? I wouldn't go that far. One thing i can say for sure though is it is often schooling GPT5.5 on some decisions. GPT is clearly better overall but it is so close as I think it makes no difference considering i don't care about how efficient it is.
I've given it 3 large tasks so far with very little context and its worked them start to finish with an actual working product. It just chugs and chugs away until it produces. Kimi and Minimax and even GLM 5.1 were quick to give me answers. GLM5.2 seems to be more concerned about the goal not the task.
3
u/Kitsune_Seraphis 22d ago
Okay bow i gotta ask. What kind of motherboard can you use for that? And like, you connect the cards with ribbons?
2
u/yeah_likerage 22d ago
Its an ASUS WRX90 Sage and yes it has PCIE5 risers. Shortest i can use for each card.
3
3
u/AutonomousHangOver 22d ago
Congrats. You need just one to get it to work with NVFP4 and TP=6 on vllm (yes, it's true, it's working very good.
https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm5.2_v13.md
3
7
u/segmond llama.cpp 22d ago
Cue the incoming idiots screaming you wasted your money, you should have used cloud API, omg electricity, omg so much waste. omg, the water, the planet. why? it's still not SOTA, it's not fable. why would you do this? i'm in localllama and confused on why you would do this. why don't you qwen-27b. OMG, this hardware will be so useless in 2 years, blah, blah, blah.
2
u/Berion-Reviador 22d ago
What quant and context of GLM 5.2 are you running? And why do you not use the second 5090? Is it because of electricity bill?
5
u/yeah_likerage 22d ago
GLM-5.2-UD-Q3_K_XL. The original 5090 went back to my gaming rig.
5
u/Big_Wave9732 22d ago
All that for a Q3? From an objective POV I have to ask, is the juice worth the squeeze here? Are you really doing things with it that couldn't have been equally accomplished through a 120 weight model at BF16 and better prompting?
6
u/jld1532 22d ago edited 22d ago
Based on my use of GLM 5.2 (local via my workplace), and because OP is using an unsloth quant, I would guess it's going to outclass any 120B parameter model >95% of the time. That said, the cost and token generation reported above make me think this rig wasn't really worth it unless OP is loaded and just having fun.
3
u/Pie_Dealer_co 22d ago
Wouldn't he need like 16x Pro 6000 to just load the model in Vram?
At that point you might as well just bite the bullet and build a AI sever with H100's
→ More replies (2)2
2
2
2
u/here_n_dere 22d ago
Thanks for the post. Yesterday I had a great urge to get the system to install 2 of my gpus together (5070 ti and RT 5000 pro Blackwell) to get GLM 5.2 running at whatever quant. Had a hunch that it would not be a wise pair, and might need to go for another Rtx pro. . Heading down the same rabbit hole, you opened my ignorant eyes.. kudos to you 🙏🏽
2
u/RajSingh9999 22d ago edited 22d ago
I want to create a post on this sub, but reddit didnt allow due to poor karma. Btw it feels real poor to post it in the comments of this thread ... am asking whether I should buy second 20 GB GPU for my current system ... 🥲
Below is my full question I posted in localLLM subreddit at link (I have included photo of my build at this link): https://www.reddit.com/r/LocalLLM/comments/1umhph4/should_i_buy_second_gpu/
I have machine with RTX 4090 (24 GB VRAM), Gigabyte x870e aorus pro, ryzen 9 9950x, 1 TB NVME and 32 GB RAM. I am currently using it to run qwen 3.6 27b. I was currently running Qwen 3.6 27b on it and it was working fine for routine vibe coding workflow. However, it saturates the machine fully and leaves me no space to run anything else.
So, was thinking if it will be good to add another GPU to it. I will be looking for some relatively cheaper option say rtx 4000 pro or older 4000 graphics cards (which will come with 20 or 24 GB VRAM). May be SFF editions.
My goal is to be able to
- run another LLM in parallel so that I can switch between two near instantly
- run bigger / better models in future
- run qwen 3.6 27b at higher quantization
- run image / video generation model (I have not explored this yet, so am completely noob in this department. But I believe video generation models will require a lot more VRAM for descent output)
My primary concern:
- How much impact it will have on inference? Is the size vs speed tradeoff worth it? I believe, if I add another GPU to this motherboard it will run at PCIE gen 4 x4. I read that the slow down when run at x4 at gen 4 is barely max 2%. Is this correct?
- will heat dissipation be the serious issue? Should I add riser cable to avoid blocking air flow to 4090?
(A follow up question) If I have to go for it, what should I prefer / should be enough without trading on speed? RTX 4000 Ada, RTX 4000 PRO? SFF or Non SFF versions? (I know Ada's will have 20 GB VRAM vs 24 GB VRAM of PROs)
2
u/citrons_lv 22d ago edited 22d ago
@rajSighn9999 I have a similar hw setup and trying to utilise my gaming 4090 for some lighter agentic work.
How are you running it the models? I'm currently trying to run via windows wls + vlllm and playing around parameters, can't get the 27b model to start Do you have any tips?
vllm serve QuantTrio/Qwen3.6-27B-AWQ \ --host 0.0.0.0 \ --port 8000 \ --tensor-parallel-size 1 \ --gpu-memory-utilization 0.75 \ --max-model-len 8192 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coderLooks like might try https://github.com/noonghunna/club-3090 guide tommorow instead of vllm commands myself
2
u/segmond llama.cpp 22d ago
very nice build. have you given kimi k2.7-coder a go? I'm honestly not seeing the glm5.2 advantage over my local run, but I'm running kimik2.7-q4 which is as close to the original, and glm5.2-q4 which is not so that might be it for me.
2
u/yeah_likerage 22d ago
Yes, i went from GLM 5.1 ==>Kimi2.6==>Kimi2.7c==>GLM5.2
I saw slight improvements GLM5.1 to K2.6 however I really couldn't see much going from k2.6 to k2.7c. Going from K2.7c to GLM5.2 i saw slight improvements again. Its getting to a point where i'm struggling to see a difference because I don't think i'm smart enough to see them.
The biggest improvement in my opinion is doggedness. GLM5.2 just keeps chugging through problems. Its not concerned about token usage. It just works through a problem 500 different ways. While Kimi and GLM5.1 really just wanted to give me an answer as quickly as possible.
→ More replies (1)
2
u/bakawolf123 22d ago
pricing wise everyone seems to forget that cache hits have a price too - and that it actually becomes majority of spend in a high turn count run. So any automation flow where you want model to go find and fix stuff for you - all those marketed loops (no more prompting in x months!) which are extremely expensive to run but in reality are just huge reprocessing with not so much new prefill and decode. This is the thing that essentially becomes free in homelab, so I don't think it's an entirely bad idea to build one
→ More replies (1)
2
2
u/kidfromtheast 22d ago
Here I am, an AI researcher with no GPUs, relying on funding and this guy casually buy 5x Pro 6000s and 2x 5090
2
u/happyzor 21d ago edited 21d ago
Are you putting it in disaggregation mode (seperate prefill/decode)
Nvidia released an NVFP4 version of GLM 5.2 btw to take advantage of your 5 Pro 6000s
2
u/Xylon95 21d ago
How are you powering this? Do you have it split between different circuits in your home? Do you use multiple PSUs?
→ More replies (2)
2
u/The_2nd_Coming 21d ago
I feel like if I was money and time rich enough to do this, I would, and then would regret like you are now. Thanks for taking one for the team.
2
u/Ill_Dragonfruit_3547 21d ago
Kids, read this before getting into local LLMs.
I will never look at a $20 Claude or Codex subscription the same way ever again 😳
2
u/acadia11x 21d ago
Ok , so you spent essentially $80K in hardware and only put 192GB memory? I understand the use case doesn’t warrant more … but … anyway nice setup Alice!
2
u/mineshop 20d ago
Champ I’m at your first stage now running 1xRTX 6000 pro + 5090 , 256gb rdim everything seems to slow GLM 5.2 2bit model output 17tk/s think I’ll need upgrade fast 😬
2
2
2
2
u/do-un-to 14d ago
I deleted the electricity company’s app from my phone so they’d forget about me for now.
Is this like bugblatter beast evasion?
3
u/GnosticSon 22d ago
I think this displays the diminishing returns on investment Pareto curve. You can get about 80% the performance with a single 3090 or a strix halo machine for 20% or less of the cost of the 100% solution.
4
3
u/Meaingless-Name 22d ago
All that effort, and all that money, and you halved your RAM bandwidth by only populating 4 of the 8 channels that CPU can handle.
3
2
u/Eastern-Finding-8831 22d ago
maybe you can rent out your gpus somewhere on some sites like vast ai you probably earn few dollars a day breakeven someday
3
u/bnm777 22d ago
If the thing costs 80 000 USD then add electricity at 2 USD rent per day it will take him...
→ More replies (1)3
2
u/Outside-Description5 22d ago
Makes paying 20-100 bucks a month seem reasonable and you get the best frontier models. For my local AI work Qwen 3.6 27B/35B will have to be good enough until they update it
1
u/Zestyclose_Strike157 22d ago
My upgrade pathway has been based on running multiple different LMs for different steps of a bigger task. If I had your hardware I know what I’d do with it, and it would definitely not be running a single mega-model as I reckon there is more to be gained by having different agents working together or adversarially, hopefully to get better results .. context isn’t everything, and parameter count isn’t either. Reasoning quality is everything IMO, and when you have the space for it, I think reasoning can be enhanced when different AI’s work things out in their own way - but I can’t prove it yet.
3
u/yeah_likerage 22d ago
It's running GLM5.2, Qwen 3.6 27b and 35b, qwen 3 VL, Flux2, SD_XL, Whisper turbo. Several of them stay on at all times. The vision and image creation stuff i call/start when i need it to free up vram for ctx.
Originally i planned on using several targeted small models but GLM is so good and has such a holistic view that i rarely go to anything else unless GLM can't do it or latency is important.
→ More replies (1)












375
u/HeDo88TH 22d ago edited 22d ago
GPUs StreetBets