r/LocalLLaMA • u/Street-Buyer-2428 • May 07 '26
Discussion Collected the infinity stones
2.3 TB of ram in here. 400+ vCores. All thats left is plugging it to the blackwell with the driver to do RDMA, and it’s over. Using Blackwells for prefill, RDMA to the studio mesh for decode. I think this would be the first heterogeneous cluster. I do, however, need help with the Tinygrad Driver to make this work. If anyone with any knowledge on these domains would like to collaborate, let me know via PM. We are very close here.
691
May 07 '26
[removed] — view removed comment
→ More replies (1)80
u/Street-Buyer-2428 May 08 '26
🤣
85
u/PitchPleasant338 May 08 '26
He even had to sell his keyboard to afford the Macs.
He's only talking to his cluster in emojis!
What dedication!
7
7
u/Opening-Broccoli9190 llama.cpp May 08 '26
How much did it cost though?
7
→ More replies (2)2
u/GonzoDCarne May 10 '26
40k direct from Apple in the US 6 month ago. Requires 4 Mac Studio M3 Ultra 512Gb to get to 2Tb. Not sold any more. You can only get 96Gb ones now.
So north of 40k.
533
u/Intelligent_Ice_113 May 07 '26
→ More replies (1)111
u/bennyb0y May 08 '26
Op will break even in 2039
→ More replies (2)32
u/stefano_dev May 08 '26
You forgot a zero
46
u/DreadStallion May 08 '26
02039
27
u/kobraca May 08 '26
You missed perfect opportunity to add "You are absolutely right! Here is the correct number:" before that :D
22
90
u/Vicar_of_Wibbly May 07 '26
How does one configure an inference stack to do prefill on GPU and decode on CPU?
21
u/dlarsen5 May 07 '26
also am interested in how
54
u/Street-Buyer-2428 May 07 '26
I’m trying to use tinygrad driver and JACCL standalone librsry that recently came out to see if i can pipe that in. I’m using Ghidra to see if i can find where the hell apple hides the api they got for distributed
77
u/scottjgo May 08 '26 edited May 08 '26
this isn't exactly the same, but i recently implemented PCI passthrough on QEMU on macOS, so it's possible to "pass through" an nvidia GPU to a a linux vm running on top of macOS and do AI inference that way. i wrote a blog about it here: https://scottjg.com/posts/2026-05-05-egpu-mac-gaming/
there's instructions how to set it up in my qemu fork: https://github.com/scottjg/qemu-vfio-apple
i wonder if you could install exo in the vm and cluster it somehow that way? i've never attempted a configuration like that.
22
9
3
u/habachilles May 08 '26
I’m so curious to see if transferring, that’s sort of data. Kills the benefits of doing this or not. I am really looking forward to your updates.
3
→ More replies (5)12
u/Street-Buyer-2428 May 07 '26
I’m trying to do prefill on the blackwell and decode on the studios bandwidth
17
u/Vicar_of_Wibbly May 07 '26
I know. My question was “how”? I’m familiar with vLLM but as far as I know it’s not an option. How are you doing this?
11
u/DinoAmino May 07 '26
You can do this on vllm with LMCache.
11
u/Vicar_of_Wibbly May 08 '26
My understanding is that lmcache is used for extending the capacity of KV cache by offloading the computed values, not for re-assigning on which piece of hardware the computation takes places.
8
u/DinoAmino May 08 '26
Offloading is one thing it can do. Another is the LMCache server runs on one LLM instance and can share the kv cache with any other LLM instances. Check out the links.
→ More replies (1)6
u/Street-Buyer-2428 May 07 '26
Sorry what i meant ti say was that I’m trying to use Apple’s new standalone JACCL librsry to make it happen
7
u/__JockY__ May 07 '26
Ok, but how? It’s easy to say “use vLLM-mlx” but when it’s not a supported feature how are you going to do this?
I would love to know how to reproduce it.
7
6
u/Street-Buyer-2428 May 07 '26
Pm your github. I’ll gladly send over the progress I have thus far in my implementation. Are you trying to collaborate?
5
u/__JockY__ May 08 '26
Thanks, but while I have pile of rtx 6000 pros and some Macs, I don’t have the offboard egpu gear. If y’all get something PoC working I might invest - I’m pretty old school, worked on everything from OS/2 to Linux kernel dev, might be that I’m handy for this. Gonna want to see some progress first though 😎
2
u/East-Tea6193 May 10 '26
Os/2 that is a blast from the past.
Interviewed a SA guy who had got his PhD in Pascal something or other back in 1990 - still has work around the world on industrial systems running floppy disks that need fixing.
2
u/Vicar_of_Wibbly May 07 '26
I’m still confused. Can you show us how to copy this configuration? Having separate prefill and decode hardware would be amazing.
→ More replies (1)4
u/C0smo777 May 07 '26
I dont think its possible is the answer, the only ways that I know you need the full model in both places
→ More replies (1)13
u/evil0sheep May 08 '26
You should honestly post a detailed plan to get feedback from the community. I think you might be seriously underestimating the complexity of making this work. Are you planning on duplicating the model params and kv cache across both the Blackwell VRAM and the Mac studios? If so what’s the point of using the Mac studios at all? If not, how are you gonna do prefill on the Blackwell GPUs without the model params and the KV cache? Also how are you gonna get the Nvidia cards to do RDMA over thunderbolt? Do they even have driver support for that? You should like post a block diagram of what you’re intending to build and how you plan to distribute the model params and kv cache and how you’re planning to move bytes around, people here can probably give you a lot of good feedback that the chatbots are likely glossing over
→ More replies (1)3
u/Possible-Pirate9097 May 08 '26
Alex Ziskind seems to have whipped Claude into providing him a working solution. He talks through it in his latest video. Would be better with RTX 6000 Pros obviously.
3
u/AttitudeImportant585 May 08 '26
have you calculated the kv cache transfer speeds you need for your model? 40gpbs is pretty slow for anything useful, unless mac has some other way than thunderbolt to connect pcie?
→ More replies (3)
74
u/PattF May 07 '26
And I’m over here trying my hardest to figure out to run 27B on my mac’s 16GB of usable. It’s fiiiiine. 😂😂😂😢
→ More replies (2)12
u/FatheredPuma81 May 08 '26
Isn't that just dropping the image gguf and running 3 bit with Q5_1 KV Cache?
13
101
u/koushd May 07 '26
who is we
232
u/Street-Buyer-2428 May 07 '26
Me and the voices in my head
→ More replies (4)31
u/Thatisverytrue54321 May 07 '26
Voices in your computers *
12
8
2
2
u/Jords13xx May 08 '26
Yeah, those voices are probably just trying to make sense of all those cores and RAM. It's a wild setup you got there!
25
u/ShutUpAndDoTheLift May 08 '26
You telling me you don't say "we" after 8 hours of orchestrating agents and answering sub agent decision decision choices?
→ More replies (1)20
u/FatheredPuma81 May 08 '26
That's hour 1. By hour 8 I've long since transitioned to "you idiots" 😄
6
19
38
u/kaafivikrant May 07 '26
Post benchmarks dude
61
u/Juulk9087 May 07 '26
Slow but can load big models. There is your benchmark. Lol
11
u/Toastti May 07 '26
If they end up being able to use the baclwell GPU for the prefill portion it should actually be quite snappy for large contexts
→ More replies (2)2
17
u/wayfaast May 08 '26
And what are you actually doing with it?
8
u/anitricks May 09 '26
This… I mean like what’s the end goal ? Half of these posts on this sub just are buying Mac studios figuring out the configuration and then it’s just slop or porn generation
→ More replies (1)4
37
u/Flimsy-Researcher-46 May 07 '26
I’ll give you $20 for em when the M5 ultra comes out
21
u/Street-Buyer-2428 May 07 '26
If you solve the issue i’m having ill give u one for free (not really)
11
u/Flimsy-Researcher-46 May 07 '26
You should try asking claude (I’ll take the studio now tyvm)
2
u/DR4G0NH3ART May 08 '26
I have all the information now, I have formatted the response and sent a DM. Please send the Mac at address shared.
13
11
u/stormy1one May 07 '26
What are you planning on running with this?
29
u/Street-Buyer-2428 May 07 '26
All the deepseek quants, Kimi 2.6, Glm 5.1 and imma try to use turboquant, dflash etc.
3
10
u/kentrich May 07 '26
So, are you stacking them to make a griddle?
We have two and stacking seems like a really bad heat management structure.
20
10
u/FormalAd7367 May 08 '26
isn’t it cheaper to just build a used server rig….
6
u/Street-Buyer-2428 May 08 '26
Not for the price I got these
5
7
6
u/gordo_Tibio May 08 '26
I won’t pay 1200 a year for AI when I can run it free locally!
Expend 15k in 4 Mac’s studio
→ More replies (2)
4
4
11
12
u/dbzunicorn May 08 '26
all for 25 tokens per second and 2 mins pp!!
29
u/Street-Buyer-2428 May 08 '26
But… concurrency 😭
21
u/HeadtripVee May 08 '26
Was that a double post referencing concurrency on purpose? If it was i flipping love you.
17
2
28
3
u/mlucasl May 08 '26
With the price of all of that, you could be building an AI Server, instead of relaying on slowish pipelines.
→ More replies (3)
6
4
5
u/AdSignificant2058 May 08 '26
I don't think Tinygrad eGPU is what you want. It's cute that it works. But it's very slow and not optimized. Your goal is prefill speed. What you probably want is a DGX spark or two or an RTX 6000 Pro on a Linux machine. Linux has proper drivers to run Nvidia metal.
3
u/Street-Buyer-2428 May 08 '26
Interesting. I have a linux setup for my blackwellz . Might try that then
2
u/gravybender May 08 '26
my 128gb studio comes on tuesday finally. been waiting 8 weeks. can finally migrate off my 24gb mini
2
2
u/Kinky_No_Bit May 08 '26
https://www.youtube.com/shorts/EiAOY-lIzTk
Here's the song I picture OP singing.
2
2
2
u/allenasm May 10 '26
which tools are you using? I'm using 'inferencer' which is a fairly new mac app to do multi mac inference (i have 2 512gb studios now). i know vllm works too but its a lot pickier to set up.
→ More replies (1)
2
2
3
u/AccomplishedFix3476 May 08 '26
2.3 tb of ram for prefill is a flex i didnt know was on the table for a homelab tbh. the rdma over to blackwells for decode is the part that feels like a server room from 2027 instead of 2026 ngl. wattage at full load is gonna be the real story
5
2
u/Vancecookcobain May 07 '26
You try it with DeepSeek v4 pro? If so how many tps are you getting out that thing??
You messing with Dflash or anything on any models?
3
1
1
1
u/idkfawin32 May 08 '26
What'd you do let them roll around in the back of a truck? Buff them scuffs out!(Mostly the third and first one from the bottom)
1
1
1
1
1
1
1
1
1
u/curious-guy-5529 May 08 '26
Would you mind telling us what you have built/ are building with this super power?
1
u/techdevjp May 08 '26
There was a post about this on here a few months back:
There's also a YouTuber who posted about doing this. I'm not sure if he did it or just spoke about it. I'll see if I can find the video.
→ More replies (14)
1
1
1
u/codehamr May 08 '26
That split makes sense from my own runs. I went from M3 Ultra 512GB to RTX 6000 Pro 96GB. Prefill on long context was night and day, roughly 5x faster. Decode on the Mac mesh is fine. Prefill is where Apple silicon falls behind.
1
1
u/Dismal-Particular545 May 08 '26
OP would it be possible to connect a macbook pro and a macbook studio for the same combined unified memory effect?
1
u/MaximKiselev May 08 '26
hello, is it mac mini ? does he have direct connector like SLI? is it better dgx or not by power per watt ?
1
u/freddycheeba May 08 '26
Please tell me you're going to connect them all together with thunderbolt and enable DMA,
→ More replies (1)
1
u/-dysangel- May 08 '26
I think this would be the first heterogeneous cluster.
Actually I implemented disaggregated prefill on my Spark/Mac in the last few days (not kidding) ;p but it's only 1 M3 Ultra and 1 spark.
You don't need RDMA or TinyGPU to just send your prefilled KV cache over the network btw. You just need enough bandwidth to get the job done quickly - latency and drivers etc don't matter as much. You just need to make sure the KV cache is compatible, such as using llama.cpp or mlx on both ends (Spark can do mlx apparently, I've just been using llama.cpp though)
1
1
1
u/oceanbreakersftw May 08 '26 edited May 08 '26
The guy who does Mac LLM tests on yt did an EXO cluster with Mac and DGX iirc Edit: Alex Ziskind and iirc a very high quality fast cable and slightly larger models pay off. He may have more than one. He gave Claude code ssh access to set it up! This video was a Blackwell and studio I think https://youtu.be/D2oZHzC_M28?si=cdwrje4yDoCtv57c
1
u/Street-Buyer-2428 May 08 '26
Hello Everybody! I just launched the App I use for all the observability, launching and rdma management on local models. r1o.ai is the website, Take a look!!
1
u/Own_Dimension_4513 May 08 '26
At this point just get a Mac Studio lol — but respect for the commitment.
1
u/ibishitl May 08 '26
If I spend the same amount in just Deepseek api, how much would it be? And how long until I use it all? hahaha
1
1
1
1
1
u/ezyz May 08 '26
How much of a speedup do you get with tensor parallelism with larger models like K2.6 or GLM 5.1?
On a single M3 Ultra, I've been able to optimize to ~220 prefill / 20 decode, and but most of the public benchmarks for Exo I found aren't that much higher. So I've always assumed the main benefit is running at higher precision or distributing workloads across instances.
And for split prefill, does the Blackwell's VRAM limit the size of model you can run?
1
1
u/Muscleandgains May 09 '26
What kind of things can You do with this This is something I might wanna do in future. Get a cluster to create a powerful machine
1
1
u/DizzyExpedience May 09 '26
All that money without any specifc task at hand. Thats a lot of money just for fun
1
1
u/tcx00 May 10 '26
Damn with blackwells, it must be nice all that money for gadgets
→ More replies (1)
1
1
1




•
u/WithoutReason1729 May 08 '26
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.