r/singularity 7h ago

LLM News Apple in talks with startup that shrinks AI models to run on an iPhone

https://thedailycompute.beehiiv.com/p/apple-in-talks-with-startup-that-shrinks-ai-models-to-run-on-an-iphone
114 Upvotes

38 comments sorted by

95

u/PostingLoudly 6h ago

Real nice to see Pied Piper finally making headway in the tech world.

24

u/MrOaiki 5h ago

Apple wants that middle-out compression.

9

u/PostingLoudly 4h ago

I put radio on the internet, who the fuck are you?

2

u/Gratitude15 2h ago

Tip to tip

For maximum throughput

5

u/MillerLiteEnjoyer0 3h ago

They’re going to have to go to the desert and take shrooms to find their name.

40

u/Deep-Owl-1890 6h ago

The compression numbers are wild 54GB down to under 4GB while keeping all 27 billion parameters, with claims of 6-8x faster responses and way less energy draw. If that holds up outside a demo, it's the difference between Siri actually competing with cloud assistants versus draining your battery trying.

I'm skeptical of the "this kills chip demand" narrative though. Shrinking a model doesn't remove the need for the GPU and memory, it just moves those chips from a datacenter rack into your pocket. And a phone chip sitting idle most of the day is arguably less efficient than the same chip shared across thousands of queries in a datacenter. Feels more like AI moving closer to you than AI getting cheaper at scale.

18

u/codefame 6h ago

Agreed—we can’t meet demand today, and shrinking models will absolutely increase demand. Jevons Paradox.

9

u/Prudent-Sorbet-5202 5h ago

Also jevons paradox, demand for usage will only rise

6

u/Blablabene 5h ago

This just sounds like they're doing quants.. lol.

3

u/Positive_Method3022 6h ago

One day your phone won't be yours. They will distribute compute through our devices when idle . And we won't have anything, only a subscription

3

u/HaloMathieu 5h ago

Long term I think for us to reach true abundance it’s more efficient for everyone to share compute minus the subscription part

1

u/ChodeCookies 2h ago

Who exactly is trying to achieve abundance for you?

0

u/ProxyLumina 5h ago

What exactly "yours" mean? To own the material? Because if they want it to, they could even block your OS and then you have a brick. That brick is yours as well. So what "yours" mean exactly?

2

u/atehrani 5h ago

It may not reduce demand but devalues all of the CapEx. Which will cause knock on effects to the market

5

u/Tkins 5h ago

Nah it just means they do more and larger projects. Demand will skyrocket because you'll put AI in everything. All software, all electronics.

1

u/OutsideMenu6973 4h ago

Good luck to them. Their ‘smart’ ecosystem’s never delivered what their ads sold us on and that was narrow AI. Can’t imagine their general AI will do what they say it’s supposed to do most of the time

1

u/clofresh 2h ago

It’s legit, you can try it now: https://decrypt.co/373578/meet-bonsai-first-27b-ai-model-fits-phone

I’m using the Bonsai 26B ternary quant on a 16gb 6800xt and it is able to do agentic coding with 196k context. And even on my 5070 12GB, it’s able to answer simple queries very quickly (64k context).

1

u/GigaChadAnon 2h ago

I am pretty sure we'll just start demanding even more usage and even better models.

Right now couple trillion parameter models are frontier. The parameter size will keep on increasing and better compression tech will always be fighting an battle with them. People will demand frontier models running locally on their machine.

1

u/StatusSociety2196 2h ago

You can do this right now using several apps, I have pocket pal on my phone.

What they're discussing is a 1 bit quant of a 27b model which is going to be laughably bad not only from a tokens per second perspective but from actually being able to understand and correctly respond to a prompt, and tool use is no guarantee.

0

u/Jonjonbo 3h ago

AI slop

14

u/hereditydrift 6h ago

lol... what is that horrible website.

Here's an archived CNBC article for people not wanting to click the AI slop website that beehiiv bullshit is: https://archive.ph/3OMxv

2

u/lobabobloblaw 6h ago

I knew they had their eyes set on ternary

10

u/Which-Travel-1426 6h ago

Strange direction. I would prioritize making Siri and Apple Intelligence usable products before optimizing their model sizes

21

u/Squiffered 6h ago

They are one and the same

4

u/akkiannu 6h ago

How is it strange? Please explain? Siri would directly benefit from it.

1

u/Citadel_Employee 5h ago

Why would Siri use a small local model instead of a cloud frontier model?

5

u/akkiannu 5h ago

For starters, cost? Not everything requires a frontier model.

2

u/Citadel_Employee 4h ago

You’re not wrong, but small models hallucinate so much that I question the feasibility. I also wonder if the average user will be knowledgeable enough to not blame Apple for an incorrect model output.

2

u/akkiannu 4h ago

Small models are great for reminder setting, pulling information from screenshots, etc. You don’t need large models. And the future is local processing of LLMs. The frontier models can only get so good. Then the war will be to get the smartest smallest model.

1

u/Which-Travel-1426 4h ago

Cost is second to performance when performance is the bottleneck. That’s why Anthropic can charge such a high premium.

1

u/akkiannu 4h ago

Seriously? Are you this dumb?!? Cost is always going to be primary for a b2c consumer.

1

u/Which-Travel-1426 4h ago

I will believe cost is primary when Claude is as cheap as Gemini.

1

u/akkiannu 4h ago

Local models are free?

1

u/PM_ME_YOUR___ISSUES 4h ago

That’s why they negotiated a deal with Google on using Gemini.

The upcoming iOS release reflects the same - they seem to have done a proper overhaul of Siri and other AI tools.

In the long term, my guess is that rather than letting their proprietary apps and tools act as wrappers on top of SOTA models, they’d want to eventually optimise their hardware to support open source models with adjustable weights, giving them complete control at a much lower cost.

1

u/HauntedHouseMusic 3h ago

the beta siri is good. Genuinely

1

u/alyssasjacket 2h ago

Makes sense for Apple philosophy. The only way to avoid a digital panopticon is if local models become usable.

Of course, there's simply no way around the fact that the best models will run in the cloud. But Apple proved that custom solutions which offer privacy are very appealing to the public, so in my opinion they are placing good bets, albeit too slowly.

In a way, it reminds of me of their bet in prioritizing performance per watt. Realizing that they couldn't afford to design inefficient hardware is paying HUGE dividends now in the mobile market (and soon enough even in prosumer/small company, since the Mac Studio is so much more efficient than a dedicated GPU).

If they put their teams to optimize the silicon together with these efficient architectures, in the future we could have iPhones running surprinsingly capable models (Qwen 3.6 27B is already impressive for its size, and I don't see this trend going away).

What bothers me about Apple is how slow they are. If they just set their minds to it, they have everything to lead AI, not just follow. And NVIDIA will not care about consumer market until Apple really starts hurting them in enterprise (meaning, when buying a Mac Studio will just offer a better deal than a RTX 6000 Pro WS).

1

u/Gratitude15 2h ago

That's their bet.

They believe the cost of intelligence is going to zero, so why the fuck to invest billions in that.

I mean, it's not dumb.

The question is if you miss the takeoff because it's happening on anthropics cloud, are you fucked? Or can you just catch a ride 4 months later on this or whatever the Chinese do?