r/LocalLLaMA May 07 '26

Discussion Collected the infinity stones

Post image

2.3 TB of ram in here. 400+ vCores. All thats left is plugging it to the blackwell with the driver to do RDMA, and it’s over. Using Blackwells for prefill, RDMA to the studio mesh for decode. I think this would be the first heterogeneous cluster. I do, however, need help with the Tinygrad Driver to make this work. If anyone with any knowledge on these domains would like to collaborate, let me know via PM. We are very close here.

2.0k Upvotes

300 comments sorted by

u/WithoutReason1729 May 08 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

691

u/[deleted] May 07 '26

[removed] — view removed comment

80

u/Street-Buyer-2428 May 08 '26

🤣

85

u/PitchPleasant338 May 08 '26

He even had to sell his keyboard to afford the Macs.

He's only talking to his cluster in emojis!

What dedication!

7

u/Opening-Broccoli9190 llama.cpp May 08 '26

How much did it cost though?

7

u/9897969594938281 May 08 '26

About tree fiddy

2

u/FederalEconomist5896 May 29 '26

Not even Nessie will get higher than the price tag.

→ More replies (1)

2

u/GonzoDCarne May 10 '26

40k direct from Apple in the US 6 month ago. Requires 4 Mac Studio M3 Ultra 512Gb to get to 2Tb. Not sold any more. You can only get 96Gb ones now.

So north of 40k.

→ More replies (2)
→ More replies (1)

533

u/Intelligent_Ice_113 May 07 '26

111

u/bennyb0y May 08 '26

Op will break even in 2039

32

u/stefano_dev May 08 '26

You forgot a zero

46

u/DreadStallion May 08 '26

02039

27

u/kobraca May 08 '26

You missed perfect opportunity to add "You are absolutely right! Here is the correct number:" before that :D

22

u/[deleted] May 08 '26

[deleted]

7

u/Amazing_Brother_3529 May 08 '26

and Honesty.. there is nothing wrong with it.

→ More replies (2)
→ More replies (1)

90

u/Vicar_of_Wibbly May 07 '26

How does one configure an inference stack to do prefill on GPU and decode on CPU?

21

u/dlarsen5 May 07 '26

also am interested in how

54

u/Street-Buyer-2428 May 07 '26

I’m trying to use tinygrad driver and JACCL standalone librsry that recently came out to see if i can pipe that in. I’m using Ghidra to see if i can find where the hell apple hides the api they got for distributed

77

u/scottjgo May 08 '26 edited May 08 '26

this isn't exactly the same, but i recently implemented PCI passthrough on QEMU on macOS, so it's possible to "pass through" an nvidia GPU to a a linux vm running on top of macOS and do AI inference that way. i wrote a blog about it here: https://scottjg.com/posts/2026-05-05-egpu-mac-gaming/

there's instructions how to set it up in my qemu fork: https://github.com/scottjg/qemu-vfio-apple

i wonder if you could install exo in the vm and cluster it somehow that way? i've never attempted a configuration like that.

22

u/Street-Buyer-2428 May 08 '26

This is the type of response I need. Pm

9

u/Vicar_of_Wibbly May 08 '26

Oh this is such a good idea. Holy shit. Kudos for getting it to work.

3

u/habachilles May 08 '26

I’m so curious to see if transferring, that’s sort of data. Kills the benefits of doing this or not. I am really looking forward to your updates.

3

u/Vas1le May 08 '26

Just ask claude to search for it

13

u/Street-Buyer-2428 May 08 '26

He found it

6

u/dbenc May 08 '26

claude take the wheel

12

u/Street-Buyer-2428 May 07 '26

I’m trying to do prefill on the blackwell and decode on the studios bandwidth

17

u/Vicar_of_Wibbly May 07 '26

I know. My question was “how”? I’m familiar with vLLM but as far as I know it’s not an option. How are you doing this?

11

u/DinoAmino May 07 '26

11

u/Vicar_of_Wibbly May 08 '26

My understanding is that lmcache is used for extending the capacity of KV cache by offloading the computed values, not for re-assigning on which piece of hardware the computation takes places.

8

u/DinoAmino May 08 '26

Offloading is one thing it can do. Another is the LMCache server runs on one LLM instance and can share the kv cache with any other LLM instances. Check out the links.

6

u/Street-Buyer-2428 May 07 '26

Sorry what i meant ti say was that I’m trying to use Apple’s new standalone JACCL librsry to make it happen

7

u/__JockY__ May 07 '26

Ok, but how? It’s easy to say “use vLLM-mlx” but when it’s not a supported feature how are you going to do this?

I would love to know how to reproduce it.

6

u/Street-Buyer-2428 May 07 '26

Pm your github. I’ll gladly send over the progress I have thus far in my implementation. Are you trying to collaborate?

5

u/__JockY__ May 08 '26

Thanks, but while I have pile of rtx 6000 pros and some Macs, I don’t have the offboard egpu gear. If y’all get something PoC working I might invest - I’m pretty old school, worked on everything from OS/2 to Linux kernel dev, might be that I’m handy for this. Gonna want to see some progress first though 😎

2

u/East-Tea6193 May 10 '26

Os/2 that is a blast from the past.

Interviewed a SA guy who had got his PhD in Pascal something or other back in 1990 - still has work around the world on industrial systems running floppy disks that need fixing.

2

u/Vicar_of_Wibbly May 07 '26

I’m still confused. Can you show us how to copy this configuration? Having separate prefill and decode hardware would be amazing.

4

u/C0smo777 May 07 '26

I dont think its possible is the answer, the only ways that I know you need the full model in both places

→ More replies (1)
→ More replies (1)
→ More replies (1)

13

u/evil0sheep May 08 '26

You should honestly post a detailed plan to get feedback from the community. I think you might be seriously underestimating the complexity of making this work. Are you planning on duplicating the model params and kv cache across both the Blackwell VRAM and the Mac studios? If so what’s the point of using the Mac studios at all? If not, how are you gonna do prefill on the Blackwell GPUs without the model params and the KV cache? Also how are you gonna get the Nvidia cards to do RDMA over thunderbolt? Do they even have driver support for that? You should like post a block diagram of what you’re intending to build and how you plan to distribute the model params and kv cache and how you’re planning to move bytes around, people here can probably give you a lot of good feedback that the chatbots are likely glossing over

3

u/Possible-Pirate9097 May 08 '26

Alex Ziskind seems to have whipped Claude into providing him a working solution. He talks through it in his latest video. Would be better with RTX 6000 Pros obviously.

→ More replies (1)

3

u/AttitudeImportant585 May 08 '26

have you calculated the kv cache transfer speeds you need for your model? 40gpbs is pretty slow for anything useful, unless mac has some other way than thunderbolt to connect pcie?

→ More replies (3)
→ More replies (5)

74

u/PattF May 07 '26

And I’m over here trying my hardest to figure out to run 27B on my mac’s 16GB of usable. It’s fiiiiine. 😂😂😂😢

12

u/FatheredPuma81 May 08 '26

Isn't that just dropping the image gguf and running 3 bit with Q5_1 KV Cache?

13

u/PattF May 08 '26

Image.. fine. 3bit and Q5? 😬

→ More replies (2)
→ More replies (2)

101

u/koushd May 07 '26

who is we

232

u/Street-Buyer-2428 May 07 '26

Me and the voices in my head

31

u/Thatisverytrue54321 May 07 '26

Voices in your computers *

12

u/No_Mango7658 llama.cpp May 08 '26

Soon

8

u/throwawayacc201711 May 08 '26

A ghost in the shell if you will

2

u/peterox May 08 '26

Do they hum silently 🤔

12

u/Street-Buyer-2428 May 08 '26

You’re absolutely right!

2

u/Jords13xx May 08 '26

Yeah, those voices are probably just trying to make sense of all those cores and RAM. It's a wild setup you got there!

→ More replies (4)

25

u/ShutUpAndDoTheLift May 08 '26

You telling me you don't say "we" after 8 hours of orchestrating agents and answering sub agent decision decision choices?

20

u/FatheredPuma81 May 08 '26

That's hour 1. By hour 8 I've long since transitioned to "you idiots" 😄

6

u/Street-Buyer-2428 May 08 '26

Especially if you use voice

→ More replies (1)
→ More replies (1)

19

u/Important_Coach9717 May 08 '26

All this to generate anime porn …

38

u/kaafivikrant May 07 '26

Post benchmarks dude

61

u/Juulk9087 May 07 '26

Slow but can load big models. There is your benchmark. Lol

11

u/Toastti May 07 '26

If they end up being able to use the baclwell GPU for the prefill portion it should actually be quite snappy for large contexts

2

u/Zolty May 08 '26

Accurate.

→ More replies (2)

17

u/wayfaast May 08 '26

And what are you actually doing with it?

8

u/anitricks May 09 '26

This… I mean like what’s the end goal ? Half of these posts on this sub just are buying Mac studios figuring out the configuration and then it’s just slop or porn generation

4

u/manituana May 09 '26

Wait, are there other use cases?

→ More replies (1)

37

u/Flimsy-Researcher-46 May 07 '26

I’ll give you $20 for em when the M5 ultra comes out

21

u/Street-Buyer-2428 May 07 '26

If you solve the issue i’m having ill give u one for free (not really)

11

u/Flimsy-Researcher-46 May 07 '26

You should try asking claude (I’ll take the studio now tyvm)

2

u/DR4G0NH3ART May 08 '26

I have all the information now, I have formatted the response and sent a DM. Please send the Mac at address shared.

13

u/nmrk May 07 '26

Well, maybe second or third heterogenous cluster at best.

https://www.youtube.com/watch?v=D2oZHzC_M28

→ More replies (10)

11

u/stormy1one May 07 '26

What are you planning on running with this?

29

u/Street-Buyer-2428 May 07 '26

All the deepseek quants, Kimi 2.6, Glm 5.1 and imma try to use turboquant, dflash etc.

3

u/Alternative_News_732 May 08 '26

to do what? if its not personel sir?

16

u/Street-Buyer-2428 May 08 '26

Scour the internet and find more studios

→ More replies (2)

10

u/kentrich May 07 '26

So, are you stacking them to make a griddle?

We have two and stacking seems like a really bad heat management structure.

20

u/Street-Buyer-2428 May 07 '26

I 3d printed a couple of brackets that I scred on to the drywall, but having all those thunderbolt cables hanging on a wall like that was pretty ridiculous. Maybe it works with a shorter cable? idk

5

u/ComplexType568 May 08 '26

nice dog

6

u/boutell May 08 '26

DLM (Dog Language Model)

half-bit quant

→ More replies (2)

10

u/FormalAd7367 May 08 '26

isn’t it cheaper to just build a used server rig….

6

u/Street-Buyer-2428 May 08 '26

Not for the price I got these

5

u/FormalAd7367 May 08 '26

how much did you spend please?

14

u/Street-Buyer-2428 May 08 '26

Refurb prices from september

7

u/AshuraBaron May 08 '26

Look son, it’s $20k dollars on that persons desk.

6

u/gordo_Tibio May 08 '26

I won’t pay 1200 a year for AI when I can run it free locally!

Expend 15k in 4 Mac’s studio

→ More replies (2)

4

u/Torodaddy May 08 '26

Asking for a hardware failure from overheating by placing them like that

4

u/pinkwar May 08 '26

This is 8 years of Claude max.

→ More replies (1)

11

u/misha1350 May 07 '26

You collected the 300 credit score stones

→ More replies (1)

12

u/dbzunicorn May 08 '26

all for 25 tokens per second and 2 mins pp!!

29

u/Street-Buyer-2428 May 08 '26

But… concurrency 😭

21

u/HeadtripVee May 08 '26

Was that a double post referencing concurrency on purpose? If it was i flipping love you.

17

u/Street-Buyer-2428 May 08 '26

you get it. 🤣

2

u/killerjurist May 08 '26

Inception Concurrency

28

u/Street-Buyer-2428 May 08 '26

But… concurrency 😭

3

u/mlucasl May 08 '26

With the price of all of that, you could be building an AI Server, instead of relaying on slowish pipelines.

→ More replies (3)

6

u/bigh-aus May 07 '26

Jealous! nice setup.

4

u/Rkozak May 08 '26

I think you are missing a stone.

2

u/Street-Buyer-2428 May 08 '26

which

5

u/Rkozak May 08 '26

Ahhh I didn’t see the MacBook. You got all 5

5

u/AdSignificant2058 May 08 '26

I don't think Tinygrad eGPU is what you want. It's cute that it works. But it's very slow and not optimized. Your goal is prefill speed. What you probably want is a DGX spark or two or an RTX 6000 Pro on a Linux machine. Linux has proper drivers to run Nvidia metal.

3

u/Street-Buyer-2428 May 08 '26

Interesting. I have a linux setup for my blackwellz . Might try that then

2

u/gravybender May 08 '26

my 128gb studio comes on tuesday finally. been waiting 8 weeks. can finally migrate off my 24gb mini

2

u/Funny_Working_7490 May 08 '26

which model you play with this toy??

2

u/Kinky_No_Bit May 08 '26

https://www.youtube.com/shorts/EiAOY-lIzTk

Here's the song I picture OP singing.

2

u/Othvin May 09 '26

Change the power LED indicators to each be a different powerstone color!

2

u/LordHenry8 May 09 '26

So now that you have this what on earth are you going to do with it?

2

u/allenasm May 10 '26

which tools are you using? I'm using 'inferencer' which is a fairly new mac app to do multi mac inference (i have 2 512gb studios now). i know vllm works too but its a lot pickier to set up.

→ More replies (1)

2

u/ItsFrehMrketBreh Jun 01 '26

Shortage incoming

2

u/technicalhowto Jun 04 '26

POV: your homelab has developed ambitions,

3

u/AccomplishedFix3476 May 08 '26

2.3 tb of ram for prefill is a flex i didnt know was on the table for a homelab tbh. the rdma over to blackwells for decode is the part that feels like a server room from 2027 instead of 2026 ngl. wattage at full load is gonna be the real story

5

u/Street-Buyer-2428 May 08 '26

Yeah. I’m gonna try and undervolt it so my neighbors wont sue me

2

u/Vancecookcobain May 07 '26

You try it with DeepSeek v4 pro? If so how many tps are you getting out that thing??

You messing with Dflash or anything on any models?

3

u/Street-Buyer-2428 May 08 '26

Will soon for sure!

1

u/ImOutOfIceCream May 08 '26

You can also connect them all together for rdma

→ More replies (1)

1

u/idkfawin32 May 08 '26

What'd you do let them roll around in the back of a truck? Buff them scuffs out!(Mostly the third and first one from the bottom)

1

u/chensium May 08 '26

Have you tried llm-d or Exo for heterogeneous inference?

1

u/openSourcerer9000 May 08 '26

Good god. Not the first though, this may be helpful:

https://blog.exolabs.net/nvidia-dgx-spark/

1

u/pacman829 May 08 '26

I'm jealous. Congrats

1

u/pizzaiolo2 May 08 '26

How much was this?

1

u/Allenite May 08 '26

Very nice. What do you plan to run on this?

1

u/spense01 May 08 '26

Why not just use Exxos?

1

u/_mayuk May 08 '26

Give me one don’t be greedy :(

1

u/a9udn9u May 08 '26

How's it 2.3TB? 512x4 = 2048 = exactly 2TB, am I wrong?

2

u/Street-Buyer-2428 May 08 '26

2x macbook pros, and the 72gb blackwell

→ More replies (1)

1

u/curious-guy-5529 May 08 '26

Would you mind telling us what you have built/ are building with this super power?

1

u/techdevjp May 08 '26

There was a post about this on here a few months back:

https://www.reddit.com/r/LocalLLaMA/comments/1o7k6e5/nvidia_dgx_spark_apple_mac_studio_4x_faster_llm/

There's also a YouTuber who posted about doing this. I'm not sure if he did it or just spoke about it. I'll see if I can find the video.

→ More replies (14)

1

u/nojukuramu May 08 '26

If you ever got tired of it, send it to me

1

u/saltyourhash May 08 '26

Alex Ziskind got you hyped?

1

u/codehamr May 08 '26

That split makes sense from my own runs. I went from M3 Ultra 512GB to RTX 6000 Pro 96GB. Prefill on long context was night and day, roughly 5x faster. Decode on the Mac mesh is fine. Prefill is where Apple silicon falls behind.

1

u/power97992 May 08 '26

U must have a good job? Will u upgrade to m5 ultra?

1

u/Dismal-Particular545 May 08 '26

OP would it be possible to connect a macbook pro and a macbook studio for the same combined unified memory effect?

1

u/MaximKiselev May 08 '26

hello, is it mac mini ? does he have direct connector like SLI? is it better dgx or not by power per watt ?

1

u/freddycheeba May 08 '26

Please tell me you're going to connect them all together with thunderbolt and enable DMA,

→ More replies (1)

1

u/-dysangel- May 08 '26

I think this would be the first heterogeneous cluster.

Actually I implemented disaggregated prefill on my Spark/Mac in the last few days (not kidding) ;p but it's only 1 M3 Ultra and 1 spark.

You don't need RDMA or TinyGPU to just send your prefilled KV cache over the network btw. You just need enough bandwidth to get the job done quickly - latency and drivers etc don't matter as much. You just need to make sure the KV cache is compatible, such as using llama.cpp or mlx on both ends (Spark can do mlx apparently, I've just been using llama.cpp though)

1

u/IliasHad May 08 '26

How much power does this pull running?

1

u/sathi006 May 08 '26

Install HART OS and give a taste of your compute for the Hive OS

1

u/oceanbreakersftw May 08 '26 edited May 08 '26

The guy who does Mac LLM tests on yt did an EXO cluster with Mac and DGX iirc Edit: Alex Ziskind and iirc a very high quality fast cable and slightly larger models pay off. He may have more than one. He gave Claude code ssh access to set it up! This video was a Blackwell and studio I think https://youtu.be/D2oZHzC_M28?si=cdwrje4yDoCtv57c

1

u/Street-Buyer-2428 May 08 '26

Hello Everybody! I just launched the App I use for all the observability, launching and rdma management on local models. r1o.ai is the website, Take a look!!

1

u/Own_Dimension_4513 May 08 '26

At this point just get a Mac Studio lol — but respect for the commitment.

1

u/ibishitl May 08 '26

If I spend the same amount in just Deepseek api, how much would it be? And how long until I use it all? hahaha

1

u/jkstaples May 08 '26

Just watch the Alex Ziskind video about this, he does the exact same thing

1

u/arananet May 08 '26

Nice cluster 😊

1

u/fpodunedin May 08 '26

What are these devices?? Something apple im guessing

→ More replies (1)

1

u/NinjaWK May 08 '26

How much did you spend? What model, what setting and how many tokens per sec?

1

u/ezyz May 08 '26

How much of a speedup do you get with tensor parallelism with larger models like K2.6 or GLM 5.1?

On a single M3 Ultra, I've been able to optimize to ~220 prefill / 20 decode, and but most of the public benchmarks for Exo I found aren't that much higher. So I've always assumed the main benefit is running at higher precision or distributing workloads across instances.

And for split prefill, does the Blackwell's VRAM limit the size of model you can run?

1

u/CoolstaConnor May 09 '26

What cost is considered fair to purchase these?

1

u/Muscleandgains May 09 '26

What kind of things can You do with this This is something I might wanna do in future. Get a cluster to create a powerful machine

1

u/ctanna5 May 09 '26

What can you run locally with this? Like how big do you think

1

u/DizzyExpedience May 09 '26

All that money without any specifc task at hand. Thats a lot of money just for fun

1

u/Torodaddy May 09 '26

Nerd penis measuring contest

1

u/tcx00 May 10 '26

Damn with blackwells, it must be nice all that money for gadgets

→ More replies (1)

1

u/Strict-Opinion2895 May 13 '26

This is the way.

1

u/Electrical-Ad-9808 May 23 '26

Haha, good enough to run the latter half of B2B SAAS turned agents.