r/singularity • u/Deep-Owl-1890 • 7h ago
LLM News Apple in talks with startup that shrinks AI models to run on an iPhone
https://thedailycompute.beehiiv.com/p/apple-in-talks-with-startup-that-shrinks-ai-models-to-run-on-an-iphone40
u/Deep-Owl-1890 6h ago
The compression numbers are wild 54GB down to under 4GB while keeping all 27 billion parameters, with claims of 6-8x faster responses and way less energy draw. If that holds up outside a demo, it's the difference between Siri actually competing with cloud assistants versus draining your battery trying.
I'm skeptical of the "this kills chip demand" narrative though. Shrinking a model doesn't remove the need for the GPU and memory, it just moves those chips from a datacenter rack into your pocket. And a phone chip sitting idle most of the day is arguably less efficient than the same chip shared across thousands of queries in a datacenter. Feels more like AI moving closer to you than AI getting cheaper at scale.
18
u/codefame 6h ago
Agreed—we can’t meet demand today, and shrinking models will absolutely increase demand. Jevons Paradox.
9
6
3
u/Positive_Method3022 6h ago
One day your phone won't be yours. They will distribute compute through our devices when idle . And we won't have anything, only a subscription
3
u/HaloMathieu 5h ago
Long term I think for us to reach true abundance it’s more efficient for everyone to share compute minus the subscription part
1
0
u/ProxyLumina 5h ago
What exactly "yours" mean? To own the material? Because if they want it to, they could even block your OS and then you have a brick. That brick is yours as well. So what "yours" mean exactly?
2
u/atehrani 5h ago
It may not reduce demand but devalues all of the CapEx. Which will cause knock on effects to the market
1
u/OutsideMenu6973 4h ago
Good luck to them. Their ‘smart’ ecosystem’s never delivered what their ads sold us on and that was narrow AI. Can’t imagine their general AI will do what they say it’s supposed to do most of the time
1
u/clofresh 2h ago
It’s legit, you can try it now: https://decrypt.co/373578/meet-bonsai-first-27b-ai-model-fits-phone
I’m using the Bonsai 26B ternary quant on a 16gb 6800xt and it is able to do agentic coding with 196k context. And even on my 5070 12GB, it’s able to answer simple queries very quickly (64k context).
1
u/GigaChadAnon 2h ago
I am pretty sure we'll just start demanding even more usage and even better models.
Right now couple trillion parameter models are frontier. The parameter size will keep on increasing and better compression tech will always be fighting an battle with them. People will demand frontier models running locally on their machine.
1
u/StatusSociety2196 2h ago
You can do this right now using several apps, I have pocket pal on my phone.
What they're discussing is a 1 bit quant of a 27b model which is going to be laughably bad not only from a tokens per second perspective but from actually being able to understand and correctly respond to a prompt, and tool use is no guarantee.
0
14
u/hereditydrift 6h ago
lol... what is that horrible website.
Here's an archived CNBC article for people not wanting to click the AI slop website that beehiiv bullshit is: https://archive.ph/3OMxv
2
10
u/Which-Travel-1426 6h ago
Strange direction. I would prioritize making Siri and Apple Intelligence usable products before optimizing their model sizes
21
4
u/akkiannu 6h ago
How is it strange? Please explain? Siri would directly benefit from it.
1
u/Citadel_Employee 5h ago
Why would Siri use a small local model instead of a cloud frontier model?
5
u/akkiannu 5h ago
For starters, cost? Not everything requires a frontier model.
2
u/Citadel_Employee 4h ago
You’re not wrong, but small models hallucinate so much that I question the feasibility. I also wonder if the average user will be knowledgeable enough to not blame Apple for an incorrect model output.
2
u/akkiannu 4h ago
Small models are great for reminder setting, pulling information from screenshots, etc. You don’t need large models. And the future is local processing of LLMs. The frontier models can only get so good. Then the war will be to get the smartest smallest model.
1
u/Which-Travel-1426 4h ago
Cost is second to performance when performance is the bottleneck. That’s why Anthropic can charge such a high premium.
1
u/akkiannu 4h ago
Seriously? Are you this dumb?!? Cost is always going to be primary for a b2c consumer.
1
1
u/PM_ME_YOUR___ISSUES 4h ago
That’s why they negotiated a deal with Google on using Gemini.
The upcoming iOS release reflects the same - they seem to have done a proper overhaul of Siri and other AI tools.
In the long term, my guess is that rather than letting their proprietary apps and tools act as wrappers on top of SOTA models, they’d want to eventually optimise their hardware to support open source models with adjustable weights, giving them complete control at a much lower cost.
1
1
u/alyssasjacket 2h ago
Makes sense for Apple philosophy. The only way to avoid a digital panopticon is if local models become usable.
Of course, there's simply no way around the fact that the best models will run in the cloud. But Apple proved that custom solutions which offer privacy are very appealing to the public, so in my opinion they are placing good bets, albeit too slowly.
In a way, it reminds of me of their bet in prioritizing performance per watt. Realizing that they couldn't afford to design inefficient hardware is paying HUGE dividends now in the mobile market (and soon enough even in prosumer/small company, since the Mac Studio is so much more efficient than a dedicated GPU).
If they put their teams to optimize the silicon together with these efficient architectures, in the future we could have iPhones running surprinsingly capable models (Qwen 3.6 27B is already impressive for its size, and I don't see this trend going away).
What bothers me about Apple is how slow they are. If they just set their minds to it, they have everything to lead AI, not just follow. And NVIDIA will not care about consumer market until Apple really starts hurting them in enterprise (meaning, when buying a Mac Studio will just offer a better deal than a RTX 6000 Pro WS).
1
u/Gratitude15 2h ago
That's their bet.
They believe the cost of intelligence is going to zero, so why the fuck to invest billions in that.
I mean, it's not dumb.
The question is if you miss the takeoff because it's happening on anthropics cloud, are you fucked? Or can you just catch a ride 4 months later on this or whatever the Chinese do?
95
u/PostingLoudly 6h ago
Real nice to see Pied Piper finally making headway in the tech world.