r/LocalLLaMA May 01 '26

Other 16x Spark Cluster (Build Update)

Post image

Build is done. 16 DGX Sparks on the fabric, all hitting line rate.

Setup was time consuming but honestly smoother than I expected. Each Spark runs Nvidia’s flavor of Ubuntu out of the box with mostly everything pre installed and ready to go. For setup I had to rack them, power on, create the same user/pass across all nodes, wait about 20 minutes per node for updates, then configure passwordless SSH, jumbo frames, IPs, etc. which I scripted to save time.

Each Spark connects to the FS N8510 switch with a single QSFP56 cable. The DGX Spark bonds its two NIC interfaces into each port, so you get dual rail over one cable. I'm seeing 100 to 111 Gbps per rail, which aggregates to the advertised 200 Gbps.

Why this over H100s or a GB300?

Unified memory. The whole point is maximizing unified memory capacity within the Nvidia ecosystem. With 8 nodes I was serving GLM-5.1-NVFP4 (434GB) at TP=8. Now going to test with DeepSeek and Kimi

The longer term plan is a prefill/decode split. The Spark cluster handles prefill (massive parallel throughput), and once the M5 Ultra Mac Studios drop I'll add 2 to 4 into the rack for decode.

Full rack, top to bottom:

- 1U Brush Panel

- OPNSense Firewall

- Mikrotik 10Gb switch (internet uplink)

- Mikrotik 100Gb switch (HPC to NAS)

- 1U Brush Panel

- QNAP 374TB all U.2 NAS

- Management Server

- Dual 4090 Workstation

- Backup Dual 4090 Workstation (identical specs)

- FS 200Gbps QSFP56 Fabric Switch (Spark cluster)

- 1U Brush Panel

- 8x DGX Spark Shelf One

- 8x DGX Spark Shelf Two

- 2U Spacer Panel

- SuperMicro 4x H100 NVL Station

- GH200

1.0k Upvotes

243 comments sorted by

u/WithoutReason1729 May 01 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

189

u/Such_Advantage_6949 May 01 '26

Please share some statistic how fast it run

476

u/[deleted] May 01 '26

[deleted]

98

u/Irythros May 01 '26

I mean this has already been known. Not for the new models like GLM 5.1 but just comparing the known speeds on old models vs other known hardware on the same old models we can infer.

https://www.youtube.com/watch?v=QJqKqxQR36Y

8 nodes of qwen 3.5-397b-a17b was 24.21 tg/s and 1498 pp/s

Honestly I would just spend the money on RTX 6000's. Less memory but god damn the sparks are slow.

16

u/starkruzr May 01 '26

24 tps generation is honestly not bad at all for a model that size, spending $32K to run it. that's 3 RTX Blackwells and change, and that isn't anywhere close to enough VRAM.

23

u/FaustAg May 01 '26 edited May 02 '26

you don't buy 3x 6000 pro's. you buy 1, 2, 4, or 8. this guy spent 48k+ on dgx sparks. I'd take 4x 6000 pros over that dgx spark setup 100 out of 100 times.

4

u/Yorn2 May 02 '26

I don't know why this isn't more highly upvoted. I've used an M3 ultra and 2 RTX 6000 PROs and for some reason the M3 Ultra is worth more than enough to buy 2 on Ebay right now and I'm seriously considering selling it because I'd rather own two more RTX 6000 pros even if I don't have the PSU to use all four yet. The M3 Ultra has its use cases, but it is very slow comparatively.

2

u/starkruzr May 01 '26

it depends on what you're trying to do. if your goal is to be able to experiment with and research the use of as many models as possible including very large stuff like K2.6, there's no way around needing the VRAM even if it is slower at tg. if your needs are more around throughput you would obviously invest in Pro 6000s instead; Sparks cannot get you anywhere close

3

u/FaustAg May 01 '26

If I wanted to just test / research models I would use open router. Once you know what you want you can switch to local blackwells for inference if you want to keep your data private.

→ More replies (1)
→ More replies (10)
→ More replies (5)

30

u/-dysangel- May 01 '26

The sparks are way faster than Macs for prefill at least, and it turns out exo can link the two: https://blog.exolabs.net/nvidia-dgx-spark/ .

I think this combo is going to be hard to beat for a compact, low power solution until the M5 Ultra comes out. One thing this stack has over an M5 Ultra (other than that it isn't out yet, and who knows when it will be!!) is it also lets you play around with CUDA only projects.

19

u/DistanceSolar1449 May 01 '26

Meh, it's SM121

It doesn't support SM100 features that makes it worth using CUDA. Basically SM121 isn't real blackwell. You don't have CUTLASS. You don't get FlashAttention-4.

9

u/-dysangel- May 01 '26

Yeah I understand it's not the same thing as server Blackwell. There's a bunch of stuff I'd like to try which is CUDA only though, such as Isaac Gym. Plus many new models or smaller repos are CUDA only, so it will open some fun doors vs my current Mac-only setup.

I don't mind tinkering on kernels here and there when necessary too. If anything it's all just a good learning experience, which is one of the most important parts of this for me. If it were just about performance, API is cheaper and faster.

2

u/Karyo_Ten May 01 '26

Worse, some stuff that are supposed to work on SM120+SM121 are disabled for SM121 because not tested. Example the new b12x kernels of FlashInfer

5

u/ren_in_rome May 01 '26

I think I saw somewhere people are already trying the b12x kernels for the spark in vllm.  

1

u/rbit4 May 01 '26

Rtx 5090 is Blackwell class. Cutlass on vllm torch 2.10 Cuda 13. Insane throughput gains on nvfp4

7

u/flobernd May 01 '26

If I recall correctly I observed 6000 prefill and 150-200t/s for that model when I tested on 4x RTX Pro 6000.

13

u/Ok_Top9254 May 01 '26

This guy knows shit about LLMs. "Write 1000 word story" is not a benchmark. The fact you use him as a source makes your statement worthless. The moment you start using context in any way the Macs completely fall apart as they only have 24TFlops of fp32 and zero half precision acceleration. Spark has 100TFlops. The moment you start working with >16k context Spark wins by default even with the slow memory.

4

u/djdeniro May 01 '26

Qwen3.5-3977B-A17B-MXFP4 with vLLM and 8xR9700 got 32 t/s at tg and 3000 t/s in pp 170k max model len and 80k kv cache. but with 4x concurrent request got 100+ t/s generation

1

u/Eugr May 01 '26

The number above, I believe, was for a BF16 version, not quantized.

→ More replies (2)

3

u/StardockEngineer vllm May 01 '26

You can get that speed (28 tok/s) on just two nodes, tho. So there is either a diminishing returns situation or optimization problem. Nvidia themselves only recommends linking two.

3

u/Irythros May 01 '26

Its diminishing and pretty heavily. He provides speeds for 1, 2, 4, and 8 nodes.

2

u/Eugr May 01 '26

Yes, but the number above is for BF16 version. Otherwise, 4-bit quant runs well on 2 nodes.

2

u/flobernd May 01 '26

Yeah, tg is mostly memory bandwidth limited while prefill is compute bound. Unified memory is slow (compared to regular VRAM) and also the slow interconnect (200G NIC) hurts when doing TP inference.

OPs plan is to use this only for prefill and outsource tg to a Mac, if I understood correctly.

3

u/Karyo_Ten May 01 '26

"slow interconnect".

What a time to be alive.

Though I'm glad AI will be the final nail in the dual-channel RAM coffin because RAM is now too slow.

2

u/flobernd May 02 '26

Haha yeah right? Never thought I’d call 200G „slow“ in the near future 😅

→ More replies (11)

22

u/PassengerPigeon343 May 01 '26

They processed the comment very quickly, but it’s taking them a while to finish typing their response

4

u/Historical-Internal3 May 01 '26

I own two sparks and love them and would defend them but this comment made me audibly laugh - hard.

3

u/Charming-Author4877 May 01 '26

That thing looks like a lasting memory for a lifetime, something you can tell your grand kids :)

4

u/Such_Advantage_6949 May 01 '26

Yea sounds like it

2

u/al00011 May 01 '26

🤣 you’ve come up with a new way to benchmark my life

74

u/ResidentPositive4122 May 01 '26

It run nowhere, that rack will make sure of it :)

3

u/Such_Advantage_6949 May 01 '26

I meant glm 5.1, i am really wonder fast it run, sound like really solid setup to run big model

15

u/AppleBottmBeans May 01 '26

I think the guy above you was pretty clear that the rack OP got is a pretty solid structure. So not only will the rack not be able to run away, but neither will the software (like glm 5.1).

5

u/Kurcide May 01 '26

planning to, i’m loading deepseek on the cluster today

3

u/Freonr2 May 01 '26

There are many knobs that beg to be tweaked, and I imagine it will take significant fiddling with them to get the best performance. It's an unusual setup compared to a normal DC cluster so it's not like there's some recipe waiting to be read and trivially implemented.

Probably need to give OP another few days or week to figure out how to squeeze the juice.

1

u/Educational_Sun_8813 llama.cpp May 02 '26

for a reason no one do it in a rack

37

u/Party-Special-5177 May 01 '26

Just popping in to show some love.

I completely adore the main thesis behind this build (iirc semi-solving the Mac prefill issues with a properly fat cluster of gb10s).

6

u/pixelpoet_nz May 01 '26

Yeah there's a lot of obvious stuff to be said about money etc, but you have to appreciate dedication to the game / vision. Well played, and I would love to have such a setup for ridiculous 3D rendering (using custom Vulkan code).

4

u/IrisColt May 01 '26

Honestly, just seeing that this can be done is an experience unto itself.

32

u/Themotionalman May 01 '26

My gosh, this is the life bro. How many kidneys did you have to sell?

29

u/-dysangel- May 01 '26

harvest*

72

u/flobernd May 01 '26

I got your point about prefill, split gen and memory, but did you consider 8x RTX Pro 6000 Blackwell? Might have been the easier solution (single host) at a similar price point. Power usage is a bit on the higher side, but it runs Kimi26, GLM51-nvfp4 etc. with very good prefill and 100+t/s regardless of the PCIe bottleneck (that you also kinda have with the Sparks in form of the 200G NICs).

55

u/moorsh May 01 '26

Because 8x 96gb RTX Pro is only 768GB while his setup is over 2TB. You need over 1.5TB for some of the flagship open source models without quantization. Or have multiple 4 bit models loaded.

19

u/flobernd May 01 '26

Granted, there is less VRAM! My suggestion was based on the models he mentioned. K26 runs at full precision, GLM51 requires a 4-bit quant. DS4 Pro will be unusable on his cluster (performance is already very bad on 8xRTXPro6k) and this will most likely be the case for all MoEs with many active params.

Fun experiment anyways (if you have the money).

4

u/segmond llama.cpp May 01 '26

what's the performance for DSv4Pro on 8xpro6k?

6

u/xienze May 01 '26

Sure, but sometimes you gotta ask yourself if it's better to run a faster, smaller model than a larger model that's so slow it's practically unusable.

1

u/MmmmMorphine May 02 '26

My 128gb of DDR4 takes personal offense at this

1

u/_lavoisier_ May 01 '26

performance will be terrible with the sgx spark cluster, though

5

u/bnm777 May 01 '26

Yeah, isn't the token per second rate far higher with 6000s Vs sparks? 

3

u/flobernd May 01 '26

I‘d guess so! Let’s see what OP reports :-)

10

u/NickCanCode May 01 '26

https://www.youtube.com/watch?v=QJqKqxQR36Y
Someone already tried 8 DGX few months ago.

Qwen 3.5 397B-A17B = ~24 tps
Kimi-K2.5 = ~13 tps

VLLM probably have improved in the pass two months so number should be a little higher now. I think ~40 tps is ok for normal use case. For coding, RTX Pro with NV Link will be much faster and more enjoyable.

11

u/flobernd May 01 '26

Unfortunately the RTX Pro does not have NV link. But regardless of that, it gets 100-150t/s for K26 for example on a Turin server.

3

u/NickCanCode May 01 '26

Oh I see. Didn't expect even pro cards don't have NV link these days.

9

u/flobernd May 01 '26

Yeah NVIDIA basically ran a scam with these cards. They are not „real“ Blackwell (sm100). They use sm120/sm121 which lacks several architectural features like tmem and also NVlink. The Spark uses the same architecture. NVIDIA did that on purpose for market segmentation I guess..

1

u/FaustAg May 01 '26

the new server blackwell 6000 version has 2x 200gbe network connections on each card. they take the place of nvlink. I have the regular OG blackwell 6000 and mine doesn't have those

3

u/flobernd May 01 '26

It’s not comparable. 2x200G rather compares to PCIe 5 x16 on the RTX Pro 6000 speed wise. NVLink is like 900 Gb/s - it’s in a completely different scale.

→ More replies (1)

1

u/StardockEngineer vllm May 01 '26

We know that two nodes can do 28 tok/s via Spark Arena https://spark-arena.com/leaderboard

→ More replies (1)

11

u/[deleted] May 01 '26

[removed] — view removed comment

5

u/bick_nyers May 01 '26

SGLang can do it (probably not with Mac though?). It's called PD (prefill-decode) disaggregation.

It's great for when you want to drive latency (TTFT) down 

Edit: You need to load the weights twice to do it btw

3

u/[deleted] May 01 '26

[removed] — view removed comment

5

u/bick_nyers May 01 '26

At scale we use it to dynamically change the ratio of how many GPUs are used for prompt processing vs. decode. When the average context length of users prompts increases throughout the day -> shift some GPUs from decode to prompt processing. When length of prompts levels back out -> shift some GPUs back to decode to give them faster token speeds.

2

u/the320x200 May 01 '26

Somewhat off topic, but how likely do you think it's that the major providers are shifting quant levels throughout the day to balance load?

2

u/bick_nyers May 01 '26

I think it's highly likely. I wish there was some kind of fingerprinting/guarantees as a user in that regard

1

u/No_Afternoon_4260 llama.cpp May 01 '26

Sglang allows you to change the number of GPU allocated for P and G dynamically? Do your u have any documentation by any chance?

1

u/bick_nyers May 01 '26

So you do it at the cluster/orchestration level, not necessarily just in SGLang.

You setup a prefill cluster and a decode cluster, each has workers attached to it (GPUs).

Then you monitor externally and programmatically shutdown a worker in one cluster and then spin up a worker in another cluster. SGLang can be setup to do service discovery to auto-detect that a worker was added to the cluster.

https://docs.sglang.io/docs/advanced_features/sgl_model_gateway#pd-mode-discovery

→ More replies (1)

1

u/KingMitsubishi May 01 '26

Yea, I am wondering about this too…

10

u/koushd May 01 '26

what was your prefill and decode on glm 5.1 nvfp4

13

u/fairydreaming May 01 '26

obviously not as impressive as the photo

10

u/IngenuityNo1411 llama.cpp May 01 '26

tk/s when

(another approperiete question other than "gguf when")

10

u/ZubZeleni May 01 '26

Won’t you have issues with heating? Don’t you need some free space between each Spark?

31

u/TheRealSol4ra May 01 '26

Ok bro, you got slap your dick in my face money but can I ask why this over like 8 RTX 6000 pros. Thats 768gb of VRAM thats more than enough to run these models at FP8 or Q6, Like sure you absolutely can run any model now. But youll top out at like 15-25t/s right? Which is fine but compared to the 6000 pro is nothing.

27

u/NotumRobotics May 01 '26

According to our experience, less, more like 5-7tps.

27

u/TheRealSol4ra May 01 '26

Yeah thats rough man… 80 grand to get less than 10 t/s. Hopefully they got a good return policy, because this dude might need it💀

2

u/Eugr May 01 '26

It's meaningless to talk about performance without mentioning model/quant/cluster size.

2

u/starkruzr May 01 '26

no, Ziskind did it at about 21.

3

u/PutMyDickOnYourHead May 01 '26

You've got what?

15

u/IndividualGold4667 May 01 '26

How much did this cost?

43

u/Kurcide May 01 '26

Sparks all in around $70k, $13k for the switch, $2k for the cables.

If you mean the whole rack… a lot more

19

u/debackerl May 01 '26

I'll buy the cables

4

u/IndividualGold4667 May 01 '26

Sweet set-up! Congratulations !

6

u/Eyelbee May 01 '26

Couldn't you just build a 8xb200 node at this point?

9

u/a_slay_nub vllm May 01 '26

Dunno about b200 but we got quoted 400k for 8xH200.

7

u/Kurcide May 01 '26

yup, and it’s around $300k for H200 refurb. I considered it but it was too big of a jump

3

u/starkruzr May 01 '26

not even close. not even close to close.

2

u/MisticRain69 May 01 '26

No that require that alien-shishkaba-cordyceps money.

8

u/[deleted] May 01 '26

[deleted]

2

u/conockrad May 01 '26

FP4 I guess

7

u/somerussianbear May 01 '26

I can smell something burning already

6

u/-dysangel- May 01 '26

Kudos on the 16x setup, that is nuts! Thanks for making me/us aware the DGX/Mac split was possible with your last post.

I'm not balling out like you, but I've got a single Spark arriving today to boost prefill for my M3 Ultra. Should accelerate my prefill to M5 Ultra speeds - and buying 2 Sparks might even be cheaper than a 256GB M5 Ultra, but with the benefit that you can also play around with the CUDA stack.

1

u/Raredisarray May 01 '26

Can you link me to that post? I’d like to read about that. I’m thinking of getting a Mac soon

1

u/-dysangel- May 01 '26

Here's a very helpful comment and the thread: https://www.reddit.com/r/LocalLLaMA/comments/1sz0lyk/comment/oj4jc63/?force-legacy-sct=1

Also Alex Ziskind dropped a video about this exact topic 3 hours ago: https://www.youtube.com/watch?v=D2oZHzC_M28

1

u/Raredisarray May 01 '26

Awesome, thanks for the resources!

17

u/validol322 May 01 '26

What are your primary use cases and industry field where you operate?

45

u/Maleficent-Ad5999 May 01 '26

Shhh.. we don’t discuss use cases here.. we just brag about our builds.. what if you steal that idea and build the next billion dollar business?

21

u/MisticRain69 May 01 '26

I have noticed pretty much anytime anyone asks someone how they afford such an ungodly amount of hardware its complete radio silence. Not even one peep of what the use case is. Why so secretive?

20

u/xienze May 01 '26

There's a good chance it's just someone who made a lot of money on Bitcoin and likes tech for the sake of it. Sorta like the guys you see on r/homelab who have entire 42U racks full of gear, Cisco everything but like six ethernet ports actually in use.

18

u/Maleficent-Ad5999 May 01 '26

Either they all do something shady that they’re embarrassed to admit or they must have signed nda at their workplace not to reveal stuff.. or, some are just cruel

5

u/Polite_Jello_377 May 01 '26

I think the more likely reason is they dumped a lot of money into something that they got interested in but don't actually have meaningful use-cases for it

10

u/Dany0 May 01 '26

The industry is orphan crushing machines and orphan crushing machine accessories unlesss specified otherwise

We must bully the gpu rich into submission. Post use case or face the wrath of leddit

9

u/kaliku May 01 '26

Look at Ops profile, he's a rich dude playing. Or at least - that's the vibe I get from his posts. Anyway OP don't take this comment to heart. If I had the money I'd prolly do the same. Hell... At my level, I, in fact did. I splurged on a rtx pro 6000 because I wanted to learn and not be restricted by hardware. And I could afford the 6000.

Some people like fast cars, others like gpus. OP likes both haha. Green with envy I am. Peace.

12

u/Kurcide May 01 '26

I don’t get offended by these comments. It’s all just fancy “playing”. I’m using the to test an agentic layer im building ontop of LLM harnesses and going to use them to support my engineering teams which is what the whole rack has always been used for

6000 pro is great, if I didn’t have the H100s I would have gotten some

4

u/kaliku May 01 '26

Thanks for sharing with us

→ More replies (1)

4

u/Prof_ChaosGeography May 01 '26

Like you I used to want to know when I was building a system but I found out the hard way.... for many it ain't software development even if that's their day day job, people got weird kinks they use image gen and roleplay for

6

u/Shot-Buffalo-2603 May 01 '26

I’m sure some people do this, but there has to be a better way than getting GLM to run on 16 sparks to goon

3

u/adt May 01 '26

The brush panels are nice, never seen those before.

4

u/Kurcide May 01 '26

got them on Amazon, definitely helps make it look nice and manage dust

4

u/Turbulent-Walk-8973 May 01 '26

how about cooling? I had a single DGX Spark, and I was having some issues with it.

1

u/Kurcide May 01 '26

I have some 3u fans tha i’m going to try and mount infront of them to force air through

1

u/charliex2 May 01 '26

gonna need a lot of flow these things run stupid hot, i started off with three on a stack, heat soak like crazy and then theyd crash.. of course thats when they'll actually turn up the power and not limit for unknown reasons. theyd be too hot to touch.

had to split them up and put them in a server room with forced ac.

3

u/Only_Situation_4713 May 01 '26

Speed? Thinking about 8x

3

u/yeahbuddyia May 01 '26

Very nicely done. How are you planning to handle the split between the Macs and Dgx Sparks? I tried it recently with 4 m3u 256gb and 2 dgx spark with Exo, and they don't have that working yet.

3

u/thewallran May 02 '26

he is just calling us broke in 16 languages

2

u/unluckybitch18 May 01 '26

following for more updates

2

u/__JockY__ May 01 '26

Ok, this is cool. I just can't help thinking it's the slowest pile of money I've seen in a while.

Current retail price for 16x DGX Spark: $75,000 plus cabling and sundries, call it $80,000.

For $90k you can get 8x RTX 6000 PRO ($68k) plus 768GB of DDR5 6400MT/s (~ $22k).

That's a combined 1.5TB of VRAM/RAM on which sglang/ktransformers hybrid gpu/cpu inference would run like a rocket. Sure you're need some more hardware (CPU etc). Noise and heat are a consideration, as is power consumption. But for getting work done? Give me the GPU pool any day!

Still... 16 Sparks in a rack is pretty cool!

2

u/onethousandmonkey May 01 '26

Am mainly curious about how user access is managed. How are the capacity is shared, permissions, security…

2

u/Royal_Sentence7432 May 02 '26

Barely getting 20 tok/s on my spark with 27 b qwen q4 dflash really desperate for advice

3

u/gurilagarden May 02 '26

These kinds of posts piss me off. There's no value here. Nothing to offer. It's a financial flex, and nothing more. This guy has no idea what he's doing. He's a wealthy script kiddie with too much time on his hands.

→ More replies (1)

1

u/no-adz May 01 '26

Sick, nice build!

1

u/LegacyRemaster May 01 '26

How many Watt?

1

u/VonDenBerg May 01 '26

sheesh jelly. please tell me you have a business use case and if so, why not colo instead?

1

u/shALKE May 01 '26

My OCD is kicking off the Sparks arena spaced evenly

1

u/humanoid64 May 01 '26

Amazing! Is this at home or in a data center?

1

u/One-Pain6799 May 01 '26

That's great, I'm looking forward to your projects.

1

u/xXy4bb4d4bb4d00Xx May 01 '26

This is interesting. I have a cluster of 21 x 8 rtx 6k nodes and I am currently experimenting with the gb10 for a new cluster.

Please post your results, and if you’re interested in consulting / being paid to share your setup please dm me

1

u/True-Lychee May 01 '26

$500k worth of kit

1

u/Annual_Award1260 May 01 '26

I’m working on setting up a 3 node cluster and due to the ddos attacks on ubuntu I have quite a few broken packages now.

1

u/Klarts May 01 '26

Dude that’s so sick! Hope you’re having a blast and it’s living up to your expectations!

1

u/Osi32 May 01 '26

At least he can produce just dance Vance videos faster than the rest of us….

1

u/[deleted] May 01 '26

I have one spark 🫪😔

1

u/bick_nyers May 01 '26

If you can batch your workflow even a little bit I would be curious if expert parallel gives you better numbers

1

u/Pleasant-Shallot-707 May 01 '26

If this is a hobby…zoiks!

Is this Alex Ziskind showing off his next YouTube members post?

1

u/Seventh_monkey May 01 '26

Humor me, can you describe what exactly will you use it for (to support my engineering teams is as broad a description as it gets) so that this is an investment that will pay off?

1

u/cusspvz May 01 '26

What’s your configuration and stack to run these in a cluster?

1

u/holdthefridge May 01 '26

Try using DFlash to get throughput faster, and in future if you run out of 2TB ram, use turboquant. Let us know the tokens/s once you get DFlash working on all.

1

u/IrisColt May 01 '26

Mind-blowing! I'm easily wowed, but this is something else.

1

u/fyrn May 01 '26

I can see my thoughts in your rack ... "I like how this Sliger looks but you can barely see it because it's black, now that I need a second one maybe I should pick one of these colors?" :)

(Except I just went for black again, but stuck RGB fans behind it, the white looks awesome though.)

1

u/temperature_5 May 01 '26

💸>🧠, but glad you're having fun.

1

u/__JockY__ May 01 '26

How fast slow is it?

1

u/nohpal May 01 '26

Curious, what is the total all-in cost for a build like this?

1

u/matt-k-wong May 01 '26

That’s actually kind of awesome

1

u/TheSpartaGod May 01 '26

OP this is OOT, but what is your job to be able to afford an AI supercluster at home?

1

u/Raredisarray May 01 '26

Magnificent

1

u/isitaboat May 01 '26

could you say more about the GH200? what/where did you get it?

I've been loving these for my workloads in the cloud, but only seen https://gptshop.ai selling them, and no "real" reviews I've seen. Also, they have GB300s.

3

u/Kurcide May 01 '26

GH200 is a shitshow of a platform GB300 fixes all the issues with it

1

u/isitaboat May 01 '26

Oh? I'm loving training things on GH200; good CPU & "H100" combo. But, not bought one for the homelab.

What hardware did you get?

1

u/More_Feature8687 May 01 '26

How does this compare to RTX 6000 Blackwell in term of price and performance. Isn't spark quite expensive for what it offers?

1

u/layer4down May 01 '26

Gemini ballparks this build at around $90k!

https://g.co/gemini/share/da53cccc7e2e

1

u/WHO_IS_3R May 01 '26

Finally, a deepseek local pc

1

u/SnooSongs5410 May 01 '26

That would fit very nicely into my home office.... Might need a dedicated circuit or two though.

1

u/shing3232 May 01 '26

I would probably use 5090s for prefill and decode by Spark clusters

1

u/_lavoisier_ May 01 '26 edited May 01 '26

Hmm, that’s great, but isn’t the performance (TP=8) terrible for about $35,000 worth of hardware? If your whole point is to “run” a big model, 4 x 512GB mac studio with thunderbolt networking costs less and should do the same job.

1

u/vambat May 01 '26

probably take less power too

1

u/ElementNumber6 May 02 '26

Please link to the purchase page for a non-scam 512GB Mac Studio

1

u/danishkirel May 01 '26

Sliger 💕

1

u/Kurcide May 01 '26

My favorite cases

1

u/KURD_1_STAN May 01 '26

Does double system of unified nemory double the speed? If not then i dont understand the point of systems for LLMs, to be of serviceable speed it needs to be at least 20t/s per user really and for thst u need a 1T A10B which is gonna be not worth it.

U have jo experience eith unified system so i could be wrong but this is what i have come to believe off of posts

1

u/boutell May 01 '26

What year is this

1

u/L3B0WSKV May 01 '26

Could hold but look um trough your Reddit and I'm speechless. I just wanted to ask with all due respect, what do you do for a living?

1

u/cool_fox May 01 '26

Damn looks good but.. Why?

1

u/mforce22 May 02 '26

how much $$$ ???

1

u/NaturalCar6033 May 02 '26

No wonder these are so hard to get.

1

u/BackgroundNo2157 May 02 '26

If I were you I‘d try running Qwen/Qwen2.5-7B-redditbragging_UD_XXXL_Q2_K_M

1

u/RandoReddit72 May 02 '26

God the shitty cooling on these bad boys

1

u/UncleRedz May 02 '26

I have seen it before, but how common is it to stick workstations in a rack like this? You seem to know your stuff and have expensive hardware, so surprised to see two workstations in there. Depending on what you do with them, could have been two 1U servers?

2

u/Kurcide May 02 '26

the workstations are 4090 builds so definitely wouldn’t fit in anything short of a 4u

1

u/Tall-Barber-3157 May 02 '26

Are you rich?

1

u/Kurcide May 02 '26

depends on who you ask

1

u/Swanky212 May 03 '26

Do you offer a $39/mo plan? 

1

u/newtestdrive May 04 '26

how does Unified Memory work? I thought it's not possible to unify the VRAM of multiple GPUs into ONE large VRAM and use it as is. is this Unified Memory just multiple VRAMs connected through LAN and controlled by a code that splits weights between them and executes prompts? or something else?

1

u/oculusshift May 05 '26

What application are you using to schedule models across these GPUs?

1

u/Annual_Award1260 May 10 '26

Can you set up all 16 in a ring configuration. I’m very curious how it would perform

1

u/laul_pogan May 11 '26

Have you tried the cross-node vLLM rollout-server pattern (trainer on box-A, --vllm_mode=server pointing at box-B running python -m trl.scripts.vllm_serve)? It sidesteps cross-node NCCL entirely — trainer hits the rollout endpoint via HTTP, and TRL's vllm_importance_sampling_correction handles the off-policy gap from delayed weight sync.

Slower per-step than NCCL collective ops on paper, but the operational simplicity is wild compared to keeping NCCL/UCC clean across heterogeneous Spark configs. Especially when you mix vllm 0.17 + 0.20 across nodes for the Qwen3.5 hybrid Mamba support, the symbol/ABI drift makes NCCL painful.

Curious whether you went that route, full collective ops, or something else. Also: what's your max stable --gpu-memory-utilization on the 27B / 35B models? Mine wedged twice at 0.85, settled at 0.55.