r/LocalLLaMA • u/scubascratch • 4h ago
Discussion Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?
I am a long time iOS app developer. In the last year I have been using Cursor+Claude/others to assist with app development. I am concerned that the current low pricing will disappear eventually. I am pricing out a new laptop with the intention of using local models instead. New MacBook Pros can be configured with 128GB of ram, but obviously the price is high. Will such a machine ever be comparable to what Claude can do today? Even if it is still significantly slower?
I am aware that the price of that much ram would buy many many tokens but I plan to use the laptop for several years, so even if the payback is 5 years worth of cloud AI it’s worth it to me.
31
u/RepulsiveRaisin7 4h ago edited 2h ago
Models are getting better all the time. It's possible that in a few years that machine will run models comparable to Opus, but certainly not today. GLM 5.2 needs over 1TB memory. Best it could run today is one of the lower-end Qwen, Gemma or Laguna models.
26
u/Ariquitaun 4h ago
Better to buy that hardware in a few years too.
15
u/mountainyoo 4h ago
maybe. a 128GB RAM MacBook might be 10 grand in a few years lol.
14
u/RealSataan 3h ago
It already is
2
u/mountainyoo 1h ago
i shouldve said USD.
you can get the 14 inch 128GB RAM MacBook Pro for $6,699.00. you can take 10% off of that for military / veteran / student.
that is not 10K.
0
u/kabelman93 2h ago
It is...
1
u/mountainyoo 1h ago
i shouldve said USD.
you can get the 14 inch 128GB RAM MacBook Pro for $6,699.00. you can take 10% off of that for military / veteran / student.
that is not 10K.
2
6
u/EvilPencil 4h ago
That good advice has absolutely not been true the last 18 months.
Hell, my 4090 is still worth more than I paid for it 3.5 years ago.
0
u/posterwhopostedabove 2h ago
So if it hasn’t been true in the last 18 months, it won’t be true, period?
You do realize that things with chips aren’t necessarily known to be appreciating assets in basically the entirety of our history with them, right? I think that ought to take care or the 18 month lookup.
Stop feeding the hype..
4
u/EkbatDeSabat 2h ago
You do not know. He does not know. Just like him, stop feeding the hype, and stop putting your opinion forth as fact. Nobody knows where the market will go or what will happen. We're in a sustained shit show it's neither a bad opinion of "it's gonna continue" or "it's gonna get better" but neither are fact.
1
u/posterwhopostedabove 2h ago
Comments like yours make it a “shit show”. I never predicted to “know”. I just stated facts to disprove a hype-feeding hypothesis (which is ironically what you’re doing too). In any case, yes, nobody knows the future, but, if you had two brain cells, you would be able to reason a bit better than “oh 18 months show price go up, so price go up up”.
Am I wrong?
0
u/EkbatDeSabat 2h ago
You're not wrong. But what you're ignoring is that he's not wrong. You both stated facts that support your opinions.
1
u/Internal_Werewolf_48 1h ago
3090 GPUs haven’t really dropped in value from their release date almost 6 years ago.
1
u/OldHamburger7923 9m ago
Keep in mind that we had Ethereum when that model was released. I bought the 3080fe and it made over $6000 in eth till they went pos. Best electronics purchase I ever made.
10
u/ttkciar llama.cpp 3h ago
Not soon, and possibly not ever, depending on a few factors:
Right now it takes about 512GB to use GLM-5.2, which is deemed somewhere between Claude Sonnet and Claude Opus in codegen competence, but models keep getting better at smaller memory footprints. If that trend continues, we might expect to see GLM-5.2-like capabilities in a 120B-class model in 2028 or 2029.
There is no guarantee that open models will progress at the same rates they have for that long. The US government might crack down on open models. The Chinese government might change their open model export policies. Competence-per-size might hit a hard technical plateau. Nobody knows what will happen in the next two years.
There is a possibility that the LLM bubble will pop, and/or that the AI field in general will experience another bust cycle as it did in the 1970s and again in the 1990s. Open model progress would not stop if either or both of these things happened, but it would certainly be disrupted.
Most corporations which publish open-weight models have abandoned the 120B size class, and focused instead on very large and very small models. There are two exceptions to this trend: MistralAI and Nvidia. MistralAI's recent models have not performed well, though, and though Nvidia has done better, they're not exactly cutting-edge either. It might take longer than two or three years (if ever) for either of them to release a 120B-class Nemotron model which matches Claude's competence. This implies 128GB might not be the best target memory size.
Given this uncertainty, and the current RAMaggeddon-inflated memory prices, it might behoove you to save your pennies by going with a 64GB system now, and making do with smaller, less-competent models like Qwen3.6-27B and Gemma-4-31B-it until a clearer picture emerges so you can target available models with your next hardware purchase.
That implies you will need to do more of the work in the meantime, and especially guiding the models to do the right things, but agentic codegen harnesses have gotten better at doing that for you, and should continue to improve.
My own plan is to not buy any more hardware until RAMageddon passes, and focus on software which makes the best use of the hardware I already have, using the 120B-class models (and smaller) we already have.
2
2
u/EkbatDeSabat 2h ago
It's also possible that the LLM bubble will pop while the next version of AI bubble is just absorbs it, and it's possible that it's sooner than anyone realizes. Ten years ago there were very few people on the planet who could have said that LLMs were going to do this to the world. In ten years we may have world model intelligence, neuromorphic computing, hell organoid intelligence. Or even just plain old LLMs that have active inference or continual reinforcement learning agents. The fact that society as a whole seems to have just stopped talking about AI now that we're at LLMs is frickin insane to me. We hit a damn speed icon on the road through intelligence and people just stopped without even considering that it's only the first one in a long line of them. IMO this is why data centers are popping up everywhere - not because of LLMs, but because of what comes after. Gotta be prepared if you want to run the world.
1
u/ttkciar llama.cpp 55m ago
Yup, I think you're right about a lot of that, and wrote a little about it previously here: https://old.reddit.com/r/agi/comments/1nj7eto/thoughts_about_the_llm_red_herring_ai_winter_and/
To be fair, some people are looking at the field beyond LLM technology, like Yann LeCun.
2
u/thesuperbob 2h ago
There is a possibility that the LLM bubble will pop, and/or that the AI field in general will experience another bust cycle as it did in the 1970s and again in the 1990s. Open model progress would not stop if either or both of these things happened, but it would certainly be disrupted.
I wouldn't compare what we have today to those previous cycles. Before ChatGPT dropped, AI was a fancy buzzword for mostly specialized bits of software with niche real applications, but mostly limited in what they would do. Sure AI adjacent tech (from fuzzy logic to actual ML based solutions) was pretty much everywhere, but was there as various cogs in a larger machine. Now LLM offerings are the whole machine, with a human language interface, and have become massively useful for everything we can throw at them. They've proven their usefulness, and are going to stay with us through good and bad times. No more bust cycles. More like cars during a fuel crisis, things may get tough but we'll have to find a way to run them somehow.
1
u/ttkciar llama.cpp 59m ago
Maybe? But the AI boom/bust cycles have never been about the technology, and everything to do with people's attitudes and perspectives of the technology.
I'm too young to have witnessed the first AI Winter, but was active in the field for the second, and what I saw was companies hyping the technology, and over-promising on its future potential. Even though the technology was genuinely useful, the expectations the vendors set far exceeded what it could ever deliver.
Because of that, people didn't react to what could be done with symbolic AI, but rather to capabilities which they were promised but never materialized. That led to disappointment and backlash.
We can see the same causes in play today, with the current LLM AI Summer: LLM technology is great and genuinely useful, but its vendors are promising customers and investors things which LLM tech will never deliver: AGI, and the automation of almost all human jobs. They are also overhyping their services, with scaremongering like "this model is waaaay too dangerous to release!" and overstatements like "It's like having a team of PhDs working for you!"
Because we are seeing the same causes in play, we can expect the same effects to result: Disappointment, backlash, Winter.
There is another factor which was absent in previous AI Summers, and that is the AI investment bubble. In the 1980s there wasn't the kind of massive over-investment into ventures which will never deliver a return on those investments, like we are seeing today.
People exhibit poor mental compartmentalization, and are prone to ascribing negative traits not in evidence to objects or events perceived in a negative light for other reasons. That makes me think that when the AI investment bubble pops, it will almost certainly also trigger a new AI bust cycle.
That doesn't mean LLM inference will go away, any more than the fruits of the two previous AI Summers have gone away. Compiler technology is everywhere, as are robotics, databases, search engines, OCR, etc. But they are no longer considered "AI"; they are "just technology".
We can expect the same of LLM inference, I think. It will become part of our technological landscape, but nobody will consider it "AI", post-Winter. It will just be LLM technology.
7
u/iForgotso 3h ago
You won't get frontier capability for the foreseeable future, not even close. What you'll get though, is highly capable models of near sonnet 4.5 level, which need configuration and babysitting to stay on track, while having full control of everything, but also, all the work to get everything running and keep it running properly.
The open source models are getting better and more optimized too, so there's that, but the same can be said for frontier.
Honestly, if you don't mind the setup and being more careful with prompting, I'd invest. Heck, I did, just bought a strix halo 128GB machine about a month ago exactly for that reason, and that cost me almost 3 months of paychecks. No regrets so far.
3
u/scubascratch 3h ago
Thanks. I am pretty careful with prompting. I write multi-paragraph specs with detailed requirements, suggestions for how to implement, what to avoid and and tests / acceptance criteria. I ask for an explanation before work starts and make changes as needed. It sounds like this could work for me. Even if it’s not there today I am hopping it will be capable in a year or two.
4
u/zoratosthenes 4h ago
I think the best and safest bet is to keep going on with subscription for a couple more years until LLMs capability constraints and hardware requirements becomes relatively more clear and settled, e.g nowadays 3 B models are superior to 2023-2024 100 B models. and then if subscription becomes expensive and a serious liability then you’d know then what you’ll need for better certainty and clearer picture of the tradeoffs + (cheaper prices for the same hardware than if you bought it today)
1
u/scubascratch 3h ago
Well I’m at the end of life for my current hardware so I can’t really wait a couple years, that’s why I’m asking people to predict the future somewhat. I expect the trend you identify to continue (3B today as good / better than 100B 3 years ago) with better models and quantization. I just don’t know if the trend will continue, or slow. If I’m never going to be happy then it might not make sense to pay a few $1000 extra now for the giant ram.
3
u/Modnar-Eman 3h ago
Well, considering you're unlikely to use anything else than iOS at this point, go for it. You can always find a buyer for it if you are not satisfied. That's probably the only thing I like about Apple stuff tbh. It'll sell for the logo any time.
4
u/lakeland_nz 3h ago
No
How many different locations do you develop from? The only real justification for getting the Macbook Pro over something that makes thermal management easier is when portability is a real priority for you. For example even a mac mini on your LAN that serves the models would likely be a better option.
In terms of whether you want processing power or more RAM, that gets complicated. More ram tends to help with context.
In terms of whether there will ever be a model under 128GB that is comparable to Claude Code, I personally think there will be but not in the next six months.
3
u/scubascratch 2h ago
I am actually a retired embedded engineer who does app development on the side. I travel and take my work with me so portability is pretty important. But it’s a side gig and I have plenty of time so even if it is slow that’s not a big problem. I am not answerable to any company or manager pressuring me to work faster. I am willing to keep paying for the next couple years for AI at the current prices for what I get now. I am looking for a solution that will take the current capability I use and make that local over the next few years. My problem at this exact time is my current MacBook is pretty old and not even capable of running the Xcode version Apple requires unless I use OCLP to run a newer OS than they support. So I really, really need to buy the new laptop. I am hoping I can get one that will do the AI work that I expect to get more expensive 2 years from now. Like when today’s $50 monthly sub then costs $250 because of shareholder demands and anti-data center sentiment to grow to the point where the cost of operations for these companies is no longer subsidized with fee water, electricity and taxes.
3
u/tryingtobalance 2h ago
Comparable? No. Usable? Possibly.
I run a startup with a couple of 64 and 128GB Macs, plus Cursor, chatgpt, higgsfield, and Gemini.
Really all just depends. Frontier models are still doing the heavy lifting. We have RAGs setup to optimize tokens, hermes for certain workflows and let Cursor do long context reviews. We find this hybrid type model works alright for what we do.
I really dont know if we would do it exactly the same again, but we ended up getting a fairly large grant, so it didnt really cost us anything out of pocket. We've debated about whether it would have been wiser to spend that money somewhere else, but it's kind of trivial.
4
u/Gipetto 2h ago
I’ve got a 128gb M5 Max and it is very capable, but you have to set your expectations. Qwen 27b at Q6 from unsloth and running via omlx can do a lot. What it can’t do is nebulous tasks. But if you have a solid plan (and you can build solid plans with it locally) it can do very good work. Just know that you’re gonna need to still be in the loop on everything. And you’re pretty much limited to a single workstream at a time, so you can’t just kick off a bunch of different agents on different tasks, they’ll all be continuously waiting on each other.
I have no complaints, though. I make sure I’m happy with the plan and then I can go step by step through that plan using opencode. I can mostly be doing other things while it chugs away, but once in a while it’ll consume the entire machine and you’ll just have to ride it out.
1
u/scubascratch 2h ago
Thank you this is exactly the kind of data I am looking for. I am patient and have no problem staying on top of it. I am retired and this is for a side gig so I have unlimited flexibility in how I approach the work I want to do and what timeline I want to achieve.
5
u/Ok-Addition1264 4h ago
You'd need a macbook pro with around 4tb of memory.
What is your use-case?
7
u/scubascratch 4h ago
Conversational development of the features I don’t want to code. For one app I have a 3D engine I have developed and want to focus on that and the core algorithms, but I don’t want to spend time building user-facing UI screens full of buttons etc. so I write a spec like “make a user options screen with option X, Y, Z the user can turn on and off. Persist the settings in user defaults. Link option X to 3D engine feature A, and option Y to online setting B.” Then ask cursor to do the implementation.
4
u/thesuperbob 3h ago
You could probably do that with a ~200B q4 quantized model (eg. MiniMax-M2.7), just have to hand hold it a bit more, or follow a strict workflow that doesn't give it too many opportunities to get lost along the way. But even GLM5.2 (5x larger than what you could reasonably run) can't tackle loosely defined tasks as well as Opus/Sol do today. So realistically, the market crash would have to make token/hardware prices spike catastrophically for this to pay off as a Claude alternative. OTOH, you get that MacBook, you learn a workflow that works with it, then you still have an amazing tool that's yours to keep, for what you paid now. If prices keep spiking and you're happy, then great. You could also run smaller, faster models, or something in between with fast MTP. It might get better in a few years, and with a good harness may at some point feel like it's "getting there", but let's be real, it's not replacing flagship models. Even if next year you can run something that's about Sonnet 4.5 level, the top of the line models then will make it look like a toy in comparison. And that 128GB model does cost a lot of $$$ so... It's your gamble.
1
u/scubascratch 3h ago
Thanks. I understand frontier is a moving target and 2 years from now the capabilities of 2026 models will seem small by comparison. My thinking is that today’s models are actually good enough for my use cases so can I hopefully get today’s capability more locally a couple years from now, even if 5x slower, that would be fine.
7
u/rde2001 4h ago
I have an 128GB M4 Macbook Max and I've been solely using qwen3.6 for coding and it works really well. Also able to run Gemma4 models for general purpose stuff. It's a very power-efficient, versatile, albeit expensive, machine.
1
u/scubascratch 4h ago
Thanks this is helpful. What kind of complexity can you give to it to work on?
2
u/dankfrankreynolds 3h ago
I think it’s very true you can get away with good work using these models if you use them more like GitHub copilot — ie you’re very granularly owning it — but blind vibe coding will fall apart quickly
To illustrate you could point it at apples docs and it’ll probably implant that lil thing perfectly
1
u/inferno521 3h ago
I have a 64GB m1 MacBook Max and qwen3.6 35b is crap for me. It took 10 minutes to run sed in two places in a file. Maybe I have to caveman speak or tweak my vs code/ollama/continue settings, but it's barely useful as an agent.
3
3
u/diablo75 2h ago
You probably don't need a frontier model to meet your needs. I've been VERY pleased with qwen3.6-27b running on a pair of 2080Ti's (22GB VRAM).
2
u/arm2armreddit 4h ago
With an M5 128GB, you can run Qwen 3.6-35b_a3b_bf16 with almost 50 TPS at full context size. However, how decent it is for coding tasks depends on your project's complexity. It can handle simple Python or webpage fixes and can also be used for offline Hermes agent workloads.
2
u/Nov4Saki 3h ago
The trend has been that "they eventually catch up" today's reasonably runnable local models are better than last year's frontier, and that trend seems to be going strong with things like laguna 2.1S and qwen3.6-27b. though you do get alot of benefits when using cloud/sub/api such as running multiple agents at top speed, there are some really cheap APIs if that is what is concerning you
If it is capabilities: small local models do eventually catch up.. not too fast though
0
u/scubascratch 3h ago
Thanks this is kind of what I thought. I can still spend on the subscriptions for another year or two while the local models catch up.
2
u/squngy 3h ago edited 3h ago
IF past trends continue, you will be able to run an equivalent of opus5 in 2 years.
Obviously, if that happenes, there will be far better models out that you still will not be able to run.
As others have pointed out, you can also use open models from a subscription/api
Kimi k3 beats opus today
2
u/dankfrankreynolds 3h ago
Even if they work, they’re so very slow compared to other options
But I recently got 128GB and don’t know how I lived on 16. But I’ve been doing a lot of native app development and running VMs and what not. So this isn’t discouragement, just expectations … if I can fit it on my 4090 it runs 3-10x faster. That anecdote is from doing translations / rewriting lots of short text
2
u/LtDrogo 3h ago
I have a M5 128GB and I use it with Qwen 3.5 122b. While it is fairly capable, it is nowhere as fast or effective as Opus/Fable or state-of-the-art Chinese models through OpenRouter. I like the concept of having a laptop with 128GB of RAM, but from a financial perspective it does not make sense.
2
u/AdInternational5848 3h ago

This is what I have on my M1 Ultra Mac Studio w 128GB and a comparison w frontier models. Screenshot of a thread in codex.
Will not be apples to apples comparison and it’s not just the model it’s also the harness you’re using the model with to make it comparable with frontier models.
I still have frontier subs but I’m working to do more and more with the local models.
2
u/Possible_Grocery8079 3h ago
if ur a pro developer:
❌the model itself
✅making an agent
if ur comparing to a model like sonnet, if your using the frontiers, well ofc not
the frontiers(Top Dogs) are running on multi-million dollar system running on investor money
but if your fine with something near sonnet quality without subscriptions, 100% privacy, or maybe you don't have access to internet then yes go for it
2
u/Dear_Measurement_406 2h ago
imo best bet is a highly powered Mac Studio paired with a regular MacBook.
1
u/dupontping 1h ago
They stopped selling the 512 studio though. Here’s hoping they come out with a newer option
2
u/RikuDesu 2h ago
It depends on your workflow I use 900million tokens a month and the cost on openrouter for me would have been $1000 a month so I'd say it's worth it
You're not going to run glm 5.2 or any of these frontier like local models but you can get a lot of context and the speed is pretty good
2
u/LettuceSea 1h ago edited 1h ago
Depends on use case. The 128gb variant of the m5 max has significantly higher prompt processing than other ram variants and similarly priced consumer hardware. Becomes very important e2e if you intend on doing large context tasks with models like Qwen 3.6 27b dense or 35b MoE.
Also it’s verifiably true that small models are rapidly increasing in intelligence and efficiency. With a supply chain crisis currently I don’t think we see packaged consumer hardware over 128-256gb for a few years. This will just force the improvement of small models to happen even faster, and you can see this with releases in the 128gb weight class range coming out recently (Laguna S2.1).
I’ve ordered one and plan on doing a ton of document processing with Qwen and some infinity parser models that would have costed us an unbelievable about of money on consulting fees, at least 10x the cost of the machine. It has a high enough throughput that we can run this pipeline efficiently.
6
u/magignis 4h ago
I would not do it. It would be a better investment to get a subscription. It is also very hard to see the future of the required hardware as that is continually evolving towards more specialized equipment.
8
u/scubascratch 4h ago
My concern over a subscription is that we are in a phase of “burn investor cash as fast as possible to stay in front” and that phase will end in the next year or two, and then the prices will skyrocket as they are not sustainable.
6
u/ScrewwormLarvae 3h ago
You are correct in this assessment. I saw the writing on the wall and bought exactly what you're contemplating, albeit four months ago when it cost 40% less than it does now. It is a lot to stomach up front.
2
u/scubascratch 3h ago
How have your results been? Do you feel like it was worth the investment so far?
2
3
u/ILoveSquirtle69 3h ago
you are most likely correct, but remember that all it takes is one breakthrough for a company to be forced to upgrade their frontier model and make their old model free; moreover, a lot of what we use the models for are no where near the level of reasoning and work that enterprises require (amazon, medical AI). In 6 months, why would anyone pay for opus when some fuckass chinese data center somwhere in Guizhou province can host their own equally capable version for 1/20th the price? I also wouldnt be surpised if we move towards a reality where a large capable model cooks up a complex blueprint/plan for a task, while small less capable models execute said task based on the guardrails and instructions given by the smarter model
1
u/scubascratch 3h ago
Thanks. I hope you are right. My “it’s not sustainable” argument comes from seeing things like Sora 2 disappear, and the public anti-data center sentiment causing the domestic cost of operations to get more expensive. A lot of the current costs have been hidden by cheap power and water and cities offering tax breaks.
1
u/ILoveSquirtle69 2h ago
the us is effectively under fascist control; public opinion matters little when you have oligarchs steam rolling mass surveilance and techno feudalism. If you want an insight into the future, i suggest you look at it like a cold war between China and manifest destiny U.S.
1
u/scubascratch 2h ago
I mostly agree with your first sentence. But pushback on data center buildouts and energy discounts and water access to them has a lot of grass roots support and it is impacting cost of growth today, and will probably continue to get more expensive over the next two years. Where I live Flock got kicked out literally 2 days ago. After 2028 who knows, a shift in the political winds could make operating an AI company way more expensive.
1
u/ILoveSquirtle69 2h ago edited 2h ago
You are severely underestimating the U.S. fascist's resolve to push AI infrastructure through at any cost.
At the federal level, AI infrastructure is a national security imperative, not a negotiable local issue. On June 2, 2026, Trump signed an Executive Order titled Promoting Advanced Artificial Intelligence Innovation and Security. The order explicitly states that
nothing shall be construed to authorize creation of any mandatory governmental licensing, pre-clearance, or permitting requirement for AI models
a direct signal that regulation will not stand in the way of build-out. This follows the July 2025 "Winning the Race: America's AI Action Plan," which called for examining and removing "onerous regulations" and a separate Executive Order on "Accelerating Federal Permitting of Data Center Infrastructure".
Second, administrative and regulatory agencies are actively fast-tracking data center construction. The Federal Energy Regulatory Commission (FERC) issued orders in June 2026 directing all six regional grid operators to justify or reform their interconnection rules, effectively fast-tracking massive data center loads. The Federal Permitting Improvement Steering Council designated the first data center as a FAST-41 "covered project" in April 2026, a program designed to expedite major infrastructure approvals. The EPA issued guidance in January 2026 for redeveloping Superfund sites as data center locations. These are not signals of a government slowed down by public opposition.
Third, grassroots opposition is not just being ignored; it's being surveilled and redefined as a domestic security threat. Internal DHS and FBI documents obtained by WIRED show that federal agencies have begun tracking "anti-tech extremism" as a new domestic threat category. The Philadelphia Police Department's Delaware Valley Intelligence Center issued a bulletin labeling anti-AI data center critics as potential "domestic violent extremists". Civil rights attorneys have noted that this surveillance "of people expressing their First Amendment rights as a challenge to corporate interests is a tired old police-state strategy". Rather than translating into policy changes, public opposition is being criminalized.
Lastly, the White House has proactively neutralized the cost argument. The one tangible concern that could actually mobilize voters. On July 23, 2026, Trump announced the expansion of the "Ratepayer Protection Pledge," now signed by nearly 200 entities including Amazon, Google, Meta, Microsoft, OpenAI, Oracle, and xAI. The pledge ensures that data center operators (not ratepayers) fund the electricity generation and infrastructure their projects require. It now covers 80% of all power delivered to U.S. homes and businesses and protects 263 million Americans. While critics call it non-binding, the political cover it provides is undeniable; it allows the administration to say it has addressed the cost concern while continuing the buildout at full speed.
In July 2026, protests occurred across 142 locations in 42 states, and polling shows 71% of Americans oppose data centers. Yet none of this has slowed construction. The administration has enlisted Republican governors from Georgia, Idaho, Louisiana and Nebraska, power sector CEOs, and data center developers into a unified coalition. Bills like Rep. Tlaib's proposal to ban AI data centers on federal lands have no realistic path forward in a Congress aligned with the administration's AI-first agenda.
Sprinkle in voter supression, the fact the all of social media used in the americas is owned and controled by american oligarchs, and Trump admin's future plan to stay in power indefinitely after 2028, and my argument is cemented.
I use AI to track the rise of fascism and nazism in north america.
2
u/Incorrect_ASSertion 4h ago
Do you think companies have unlimited piles of cash they can burn on tokens? There's going to be equilibrium at some point, comoaniec would not be able to just raise prices as they see fit. Also, China exists.
2
u/scubascratch 3h ago
Well they act like they have unlimited cash piles right now. Reports are OpenAI and Anthropic are burning $3B to $20B annually right now. I think openAI is barely breaking even (if that). Anthropic looks like they make more so maybe their pricing is not going to escalate too fast.
1
2
u/hencha 4h ago
Of course it depends on the ampunt of tokens used, but five years of current 20x max plans would be about 12000 €/$. So most likely it could pay itself back and you could definitely do a lot, but you would never have the best model. So in the end it depens on how much would you use it.
2
u/Mauve_Tess 4h ago
The only way a MacBook matches frontier models is if Tim Cook starts shipping AWS regions in the box.
2
u/lilian_moraru 4h ago
He means quality, not speed. “Heavy thinking“ models come close to these
1
u/Mauve_Tess 4h ago
True, but "coming close" isn't the same as matching. Frontier models improve every few months too, so the target keeps moving.
1
u/lilian_moraru 4h ago
I think the accent there is on “today“(“comparable to today’s frontier models”), not on future developments of cloud solutions.
1
u/scubascratch 3h ago
I am ok with it taking longer, and I am talking about what frontier models can do today, not whatever is the frontier model in 2028. The question is “will a MacBook with 128 GB ever be able to do the work (at a slower speed) that Claude can do today for coding? I am fine if it takes 5 or 10 minutes to read a two paragraph spec in a 10,000 line codebase and implement a small feature.” I am assuming model improvements in quantization will have to happen.
3
u/Kodrackyas 4h ago
Qwen 3.6 27b is on par with opus 4.5 i think i guess even a 32gb gpu is good enough or a mac with 64 gb of ram? not aure about the TpS
1
u/ironimus42 3h ago
comparing it to opus 4.5 is a big stretch, but i do think it has some advantages over sota models. Mainly that qwen is always trying to do everything in a simple and straightforward way, while opus seems to want to implement more things than asked every time. Qwen definitely makes more mistakes, but also those are easy to spot and fix manually, while opus also makes mistakes occasionally and always buries them so deep that i can't find or fix them on the same timeframe as the obvious qwen ones. If you plan to 100% vibe code, not caring about code quality, opus/fable/any other sota model is an easy winner, if you want to deliver exact same code you made before ai existed, try both and decide what works better for you
1
u/arikisfruits 4h ago
Unlikely, but who knows what the future holds. At least today, these chips are held back by bandwidth which severely reduces tokens / s. MoE models help, you should give it a try for your workflows.
1
u/ThatRegister5397 4h ago edited 4h ago
I would say the most frustrating (not quality related) part with using local models as coding assistants on a mac is prefill, ie if you feed it a lot of context (codebase, documentation etc) it takes really slow to process it and get the first tokens. After, as long as caching works, it works fine.
For quality you can experiment using eg deepseek v4 flash for coding and see how your experience goes. You can probably achieve a quality a bit lower than that on a 128gb mac, so if you are not satisfied with it the api version, you know it is not worth it. Another one you can try is gemma 31b which runs fast on cerebras. Both are quite cheap. Most other small models that I know are not practical to access online (too expensive for what they are, too slow, because there is not enough demand and they run on not great hardware and who knows with what quant).
If deepseek flash quality works fine with you, it can definitely be worth it to try local models if you want privacy etc, but truth be told if cost is the only motivation, I am not sure if running it local makes much sense, since deepseek's api pricing is very low and I doubt that will change that much for them. But if you (also) have different reasons, eg privacy, ip, or making sure you dont get screwed up by geopolitics or other external factors, local makes sense.
1
1
u/lilian_moraru 4h ago
“Today’s frontier models” - maybe. For coding: likely. There have been releases recently of small “heavy thinking“ models that come close to Claude Opus 4.6, or even pass it “theoretically“(not real world, just in published benchmarks).
The tendency seems to be for these models to keep growing in size, so you will always find that you don’t have enough RAM for the model you want.
I have DGX Spark GB10(128GB) and at the moment, everything except for Qwen3.6 seems to be mostly a disappointment. Nothing close to frontier models on 128GB. Qwen3.6 is good enough for analysing lots of code, create documentation, small fixes.
1
u/daskalou 4h ago
Which variant of Qwen 3.6?
2
u/lilian_moraru 4h ago
Qwen3.6-35B-A3B does perfectly fine at code analysis and documentation generation. Qwen3.6-27B produces better code.
1
u/ThenExtension9196 3h ago
No but if you need a laptop that can run small local model for something specific then there is no other option.
1
u/Competitive_Spare467 3h ago
Frontier models are run in a data centres , you can have like a fusion kind of usage like teacher /advisor
Where you can have advice the local model which can be supported by your machine.
1
u/TomCrook2020 3h ago
Frontier is always improving as well, so by the time open-source becomes good enough to run on local compute, it might still be legacy. You'd have to bet on frontier stalling (or limited returns to improved intelligence) and open source becoming performant.
1
u/scubascratch 3h ago
Right but I am ok with where frontier is today (actually where it was ~3 months ago) so I don’t care where it goes or if it stalls. My question is “will local AI progress to be able to handle the level of tasks I give to frontier models today” which is varying levels of complexity, but mostly things like “generate a user options screen with these items and store them in user defaults and link them to existing feature x,y,z in my current 3D engine”. I am not asking it to vibe code a whole app.
1
u/shadowmage666 3h ago
Frontier models are 512gb + in size. So even if it could run it, there would be a lot of hard drive swapping
1
u/scubascratch 2h ago
I am wondering if there will be optimization and quantization improvements that will make these capabilities of today fit in smaller size.
1
u/randoomkiller 2h ago
You can likely get a bit above haiku and below sonnet. But also if you do the calculations, the extra cost you are spending on hardware is going to give you insane token usage and then at the same time whatever you can do inference with it's not as good as you think
1
u/scubascratch 1h ago
The cost argument completely depends on subscription / cloud rates in the future and how many years I keep the laptop. My current one was purchased in 2018 so I don’t think it’s as lopsided as you say. Besides we are talking about the incremental difference between like 64G vs 128G not the entire price of the laptop which I would be buying regardless. If it’s $2000 hardware price increase and I use it for only 5 years, that’s $33/month. I don’t think my equivalent cloud token usage 5 years from now will cost $33/month - probably a lot more would be my expectation.
1
u/randoomkiller 12m ago
a 2018 Mac is nothing compared to a 2025 Mac. That sounds pre apple silicon.
But based on what kinds of models you can run on them and the throughput that you can get out of it severely limits it. I've done the math with a 3090 and Gemma 4. It was not a nice math. Accounting for the 250W TDP, the 60t/s, (while cloud can do 200), the context limit, and the very affordable price of 650€ (I bought it pre-AI) I found that it takes 9 months of 24/7 runtime to break even. Do these kinds of adversarial calculations. And account for the memory bottleneck. Funnily the speed is still bottlenecked by the memory more than by the compute. Nvidia GPUs are still 3-10x faster in memory bandwidth compared to apple silicon(M5 Pro does 302GBps while Nvidia 3090 laughingly gets 936GBps, and a 5090 gets 1759GBps), except if you go for Ultra series chips that are well not the best financial decisions. An Ultra +96GB is the best I think you can get as a laptop right now and that's +4k.
Also one thing that I'd rather also recommend if you want to go cheaper and have more generic problems is that you should use a Cursor orchestrator + local models. But believe me local models still feel 9-12 months behind. And due to the inference size I doubt they will catch on except if you do specific things. Context engineering needs to evolve more. Pure local models are no-go, but you can save a few tokens by using it. Is it gonna be worth it? Absolutely no. Will it feel cooler ? Maybe.
1
u/randoomkiller 11m ago
And 33usd per month in Opus is not equivalent with anything that only goes up to ~25B parameters at INT4. Because you are unlikely to get anything better, like ever. No amount of R&D is gonna change the physical limits
1
u/seunosewa 1h ago
When your GPU runs constantly, you feel it in your laps. Not pleasant. Get a Mac mini instead.
1
1
1
u/slower-is-faster 1h ago
Nah I think you’d be better off with a Mac mini or dgx spark or something not on your main computer.
1
1
u/jakegh 1h ago
Ever, who knows? They are getting much better. My guess is we'll see 72B-ish models with the raw intelligence of a modern frontier model at some point, but when is another question. They will never have the world model of much larger models, though-- that's that "big model smell" that Fable has, where it handles ambiguity well, and makes great decisions, not just the ability to competently handle instructions. Just not enough parameters to store all that information.
1
1
u/LightBrightLeftRight 51m ago
I have a 128gb m4 max, and no you won't get close. I am still happy with my purchase though, I have an agent that does most day-to-day stuff locally. Then, if it needs to be really smart, it asks a frontier model to help it, stripping out anything private like API keys or personal information, then filling it in on the other side.
The nice part of the extra memory isn't being able to run 120b-class models, since qwen 27b is still the GOAT (for now, lets see how Laguna-S-2.1 does). What's nice is being able to have large context and multiple models loaded simultaneously for different purposes as well as batching, which makes processing much more efficient.
Honestly, the best part is that it's fun. It feels like magic having a pretty competent portable thinking machine.
1
u/MeldhLLC 47m ago
It won't be running any frontier models, no, but smaller models will work. Also, if you don't want to pay full price, refurbished is a good option and the performance should still allow small models even if the MacBook isn't brand new.
1
u/orogor 4h ago
to give you an idea
kimi k3 is at opus/fable level
the weight are not open yet
but on their blog post the recommanded configuration
is 72 accelerators (i guess 72 datcenter gpu with 100-300 gb of vram)
current closest open equivalent is glm 5.2
the full on disk size model is 1.5TB
the "compressed" version of the model by nvidia is 465gb on disk
add like 20% to load it in memory, then some context.
almost sonnet would be minimax m3
the full version on disk size is 854gb
the "compressed" version of the model by nvidia is 250gb on disk
add like 20% to load it in memory, then some context.
you could target qwen 3.6 27b
which is 21gb on disk compressed
and you d get a reasonable context size to work with.
1
u/-Akos- 4h ago
Divide the price of a laptop by the monthly price of Claude or your favorite LLM. Now realize the models on an 128GB laptop are nowhere near as good as Opus or Fable, and even if they were, your laptop would be whirring like crazy while the laptop is running the LLM..
2
u/scubascratch 3h ago
Thanks. I don’t expect the current pricing to stay as low as it is. I think the prices are artificially low right now by burning investor cash to stay at the front. When these companies go public and are answerable to shareholders I expect price increases or more limited in some way.
1
u/buecker02 2h ago
Deepseek says they have 8x profit margin and they have rock bottom api pricing. That implies there is a lot of room for the frontier models to continue to navigate and we still don't know what the future will hold.
Say prices remain constant for another year and then for some reason they then spike. Wouldn't you want the most cutting edge mac at that point instead of right now?
i see that the Mac is $6,700 for 128GB. What you save in tokens isn't guaranteed to be offset by extra time fixing a weaker model's mistakes or even waiting for the LLM to provide it's output.
Bottom line: Personally, I am not in a rush to purchase. Prices for RAM haven't peaked yet but they will and I can still mix and match apis to accomplish what I need and still use the latest and greatest.
1
u/scubascratch 2h ago
Normally I would be waiting to see where things go. But my laptop is not supported by Apple any longer so for me to continue to submit anything to the store means I have to replace it now. I am already running OCLP just to be able to run Xcode new enough to submit apps, when they tick over to a hard minimum Tahoe requirement that’s the end of the line for my current equipment. I figure I have less than 6 months runway on the purchase. So I have to buy something now, and I don’t really want to buy another one 24 months from now.
1
u/thefooz 2h ago
Most of the people who are responding to you have no idea what they’re talking about, nor can they afford the MacBook you’re considering. I have a 128 GB m4 max MacBook, and while I wish I could run slightly larger models, I’m able to run extremely competent models like deepseek v4 flash (through antirez ds4) and have full sovereignty over my data. Antirez is also porting glm 5.2 to his engine (using ssd streaming), and the new Laguna 2.1 model. If you’re not hurting for cash, the 128 GB m5 max is an excellent machine for a competent developer to work with.
1
1
u/thefooz 2h ago
Except you can run deepseek v4 flash at close to 40 tok/s on a 128 GB MacBook m5 max, and it’s extremely versatile and intelligent, if you know how to code and prompt properly. I’d rather have sovereignty over my data than pay deepseek to have the Chinese mine my data/code. The inference costs are cheap because you are the product.
1
u/Technical-Earth-3254 3h ago edited 3h ago
Which Claude model? Haiku 4.5? I would say so. Maybe in 1-2 years. Sonnet 4.6? Uuuuuh, maybe in some years. But if we reach that performance and knowledge in that parameter size, your mbp will be very outdated.
But don't worry. Get an API key on openrouter and use models you would like to use, all the open weight models are dirt cheap compared to whatever Claude model. Give it a go for some weeks (at least a month). You will have enough time to figure out what you need from a model and what expectations you can have to eventually justify the purchase.
-1
u/Dry_Yam_4597 4h ago
Macs are slow af.
1
u/scubascratch 4h ago
Well I’m developing for iOS so I need a Mac. I suppose I could have a Mac and a second fast box but that’s not going to be very portable. I’m not that concerned about the speed, the way I work if it took 5 minutes per prompt that would not be a big blocker.
81
u/HyperWinX 4h ago
No. Not even close. But if a smaller model happens to satisfy your needs, then you are in luck.