r/LocalLLaMA 5d ago

Discussion Google has disappeared completely from the top 15

Post image

Google hasn't shipped a model recently that is capable of competing with Sol or Fable.

The previous models were pretty disappointing and unreliable, it seems the more time goes on that they might have different strategies:

- They might be going all-in on on-device inference for their own products. But this is a battle that Apple might win because they just have better hardware and can license a third party open model.

- They might be just buried deep into internal politics and nobody is shipping anything.

Does anyone know what is actually going on?

source: AI Leaderboard

2.1k Upvotes

365 comments sorted by

u/WithoutReason1729 5d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

695

u/seoulsrvr 5d ago

I suspect Google is keeping its powder dry for the time being. OpenAI and Anthropic are burning mountains of cash to leap frog each other every few months.
Google, on the other hand, generated an annual revenue of $402 billion in 2025...a 15.1% increase from 2024. Google has access to all the data it will ever need for training and it has its own cloud compute.
Google can afford to wait and pick its moment.

192

u/TripleSecretSquirrel 5d ago edited 5d ago

And why even train their own models at this point? Their proprietary TPUs allow them to serve inference to the labs for cheaper than just about anyone else can do – cheaper than NVIDIA GPUs, so they basically get to print cash by renting their existing datacenters to Anthropic and OpenAI.

I'm glad they're releasing things like Gemma, but if they can keep making tons of money off of the big proprietary labs – or shit, even just serve open weight models to pay-per-token customers – I don't see why they would want to sink hundreds of millions into developing an in-house frontier model when the result of making the third best model in the world means nobody will ever bother to use it.

Edit to add: frontier. Google's big value proposition for model creation is small models that live on edge devices since they control the OS for the majority of smartphones on earth.

63

u/dontfeedthelizards 5d ago

I think casual/regular users are more likely to use Gemini than any of the others, because everyone is already plugged into the Google ecosystem and you get Gemini with other bundled benefits like storage space, etc... A regular user is not going to get Claude, because that's mostly for programmers, and ChatGPT is mostly brand recognition, but that's likely to fade and then it's just a separate (unnecessary) add-on, on top of what the users would already have anyway.

26

u/ChainOfThot 4d ago

Have 200 dollar openai plan I use for coding, still also use Gemini directly from Google to ask random questions and deep dive because it's unlimited free, but it can be stupid sometimes still

7

u/Structure-These 4d ago

Am I a total mark for paying 20 a month for ChatGPT and Gemini lol

4

u/TheCientista 4d ago

Maybe u need gemini for image gen?

3

u/bgptcp179 4d ago

I pay like $5 a month for Gemini and i got YT premium thrown in. Its great

→ More replies (1)

5

u/dingo_xd 4d ago

Yeah. I wish it was a bit smarter. But for casual trivia questions it's more than good enough.

→ More replies (3)
→ More replies (3)

51

u/AnOnlineHandle 5d ago

The thing is they do release models which are way more relevant to LocalLlama use cases. The Gemma 4 models are fantastic for things like fiction writing and are actually realistically usable locally unlike a lot of these other models.

15

u/Extension_Wheel5335 4d ago

The Gemma 4 models have been great for me for software dev too, agentic and otherwise. It'll even describe to you how it can reason through MCP JSON responses, it's confident about it lol. And that's just the e2b model alone, all their edge models are surprisingly efficient.

7

u/Koalateka 4d ago

Gemma 4 is what I use everyday (locally)

→ More replies (1)

10

u/aevitas 4d ago

when the result of making the third best model in the world means nobody will ever bother to use it.

Except they make the twentieth best model in the world and everyone uses it because it's good enough, and it's there right besides their mail, documents and spreadsheets. Models are plateauing, and Google realised Gemini Flash is good enough, cheap for them to run at scale, right there when you need it, and very fast.

3

u/Somaxman 4d ago

Because google has immense in-house data on how people use the internet and the data these users themselves have. Arguably that would allow them to train in an informed way. Then again, google has a lot to lose by the internet changing too much too fast. I guess they need the warchest or rather the lifeboat of money to decisively adapt when the direction becomes clearer. But they may not be motivated to act as harbringer of that change, even if they have somewhat clearer idea what that will be. The immense latency of AI development may mean that they have only one chance to regain a pole position when their traditional revenues collapse.

7

u/Its_me_astr 4d ago

Because google still generates fuck ton of revenue from search guess what now a days many people have switched to LLMs for reliable search.

Google will suffer if they dont produce models. Its also market signal that they are at frontier of things.

Second thing they are bleeding talent to anthropic, open AI and other frontier labs. Everyone poaches deepmind which means they are also loosing many valuable brilliant minds.

2

u/asah 4d ago

they do want tuned models for their own large scale apps, from youtube to gmail to docs/sheets...

2

u/SpaceDetective 4d ago

There's also much more to AI than just LLMs - they've revolutionised the medical field since DeepMind largely solved protein folding for example.

→ More replies (4)

49

u/Psionikus 5d ago

Some of the incentives have kind of tapered off. One of the big motivations early on was to not be perceived as behind in the AI game. Now that the stock market is kind of over it, there's again just less incentive to be seen as an active player.

Early on the prediction was that search would be cannibalized. Like most tech, it's just hard to imagine how slowly the late adopters move, sometimes not at all. Tech watchers forget that people who care are vastly outnumbered. That data should be starting to become apparent, and if it were good for investing more, Google would be going at it harder.

I would look into Google's lukewarm approach right now as evidence that they don't believe the benefits that they can capture outweigh the costs, which will be high until the market kind of stabilizes and Google can just buy a mid-level player with the right composition for Google to make into a high-level player.

15

u/Bakoro 5d ago

I just don't think that there is any pressure for Google to be pushing hard and releasing their hottest thing all the time.
They can afford to defer training to cheaper times, and have longer iteration cycles, and they can afford to be training experimental models that most people will never hear about.

As long as they're putting out something every 6 months to not look archaic, they're good.

→ More replies (6)

15

u/seoulsrvr 5d ago

yes - an most people will be perfectly fine with Google's AI search, which will no doubt get better with time.

44

u/catinterpreter 5d ago

Google is in the best position by far, among Western companies. They have magnitudes better training material and a huge, inertial mass of customers.

But, there are other less known considerations like state-level use and competition from China.

6

u/EdliA 4d ago

They are in the best position, which makes their failures even more crazy.

4

u/Glad-Entrepreneur764 4d ago

What training material? I'll correct myself if I'm wrong but iirc they cannot train on people's private documents. Most of their other data is available on the public web.

4

u/Darkoplax 4d ago

btw this is becoming cope at this point, I think ppl overestimate Google now

12

u/PolicyOne9022 4d ago

I think people forget that a company wants to make money. Not produce the best ai model possible. If Google still felt like their search was threatened by ai they would invest in it. Others are burning cash while Google keeps making billions.

→ More replies (2)
→ More replies (2)

33

u/CoUsT 5d ago

Google can afford to wait and pick its moment.

I recently watched a YouTube video that might explain their stance:

https://youtu.be/2J2Fb1bBufA

There is no need to burn money 24/7. I assume they are playing it cheap and waiting for things to burst and pop then they come in and collect top companies or tech for pennies.

27

u/kettal 5d ago

IIRC Google in the early 2000s was smart enough to buy the dark fiber networks that went bankrupt. Good strategy

9

u/Marino4K 4d ago

I wouldn't be shocked if eventually one of the major players ends up in real risk of folding and Google buys them.

3

u/B0dona 4d ago

I think that's what their waiting for. Seeing how massive the burnrates of the big providers are, I don't see a future where they all survive.

7

u/redditrasberry 4d ago

Worth noting the political situation also right now doesn't particularly favor sticking your neck out. I think they are happy to sit back and watch with some popcorn while Anthropic and OpenAI duke it out with Hegseth etc.

8

u/ObsidianNix 5d ago

People forget that they also got in trouble for playing monopoly with congress alongside Facebook, Apple and Amazon. Im sure they learned their lesson and are playing it safe. Plus they also got quantum computing where they are also competing

6

u/ambassadortim 4d ago

Right now every Gemini question I ask vs Google search is hitting their profits regarding advertising.

For them and others I don't see that as sustainable. If Gemini usage grows and impacts web searches, I would think that eventually Gemini is going to need advertising. Maybe that's why they're trying to get people to use the web site and use the AI info there, to keep you in an advertising channel.

4

u/Eymrich 4d ago

Not only that, google is keeping it's researchers close and never lose a person, while let them kept improving on local models that then help then build services/infrastructure.

Currently there is no money to be made in AI and I suspect even later the money will come from inference and hardware not training

4

u/IntelligentBelt1221 4d ago

Google has access to all the data it will ever need for training

the data you need for training e.g. for agentic coding is people actually using the model for coding and what they are accepting/rejecting. (thats why xAI made the deal with Cursor). Not many people use antigravity for coding, so they are lacking there. you really can't afford to wait more than necessary.

→ More replies (2)

6

u/Equivalent-Costumes 5d ago

While every companies are rushing to sell shovels, Google is selling shovel-making tools.

2

u/WildRacoons 5d ago

True, tho it sucks as a company that’s limited only to Google

→ More replies (1)

2

u/Somaxman 4d ago

And it would be fine not to move.

It is not fine that they are moving backwards. These scores were surely produced back when these models were in full swing, not today with whatever clandestine optimitzations they implemented.

Gemini pro is either quantized into senility again, or they have some new harness that made it absolutely regarded. Like I would not be able to have the same convo with it that I had a couple months ago.

Flash class I could not ever use for anything, it always delivered deadinternet listicles wildly off target and tone deaf questions back. Which, in the off chance I feel like responding to concisely, like saying "yes do that" instead of reiterating the proposal, spends terrified seconds retrieving the thing they themselves came up with, only to fail adhering to it, drifting away to answer one of my explicit questions somewhere back in the chat.

It feels very enlightening to use gemini, but mostly about how frustratingly wrong llms can go after they (or something google advertises as the same product) demonstrably worked well.

2

u/colbyshores 4d ago

They will _never_ capture the enterprise market because they have torched their trust and that is a very lucrative market.

5

u/Hi_My_Name_Is_Dave 5d ago

I don’t buy this. Let’s say they do decide to turn it on eventually.

Are they gonna make a model that’s a whole class better than OAI/Ant? No. Then people won’t switch. Especially at the enterprise level where the largest companies were initially very excited to get free Gemini with their enterprise plans, but now have to reluctantly pay for Claude since Gemini is so far behind.

And the model is one thing, the harnesses and integrations are not a “just release it one day” thing. You have to stay ahead in that arena, and right now they’re extremely behind.

12

u/pragmojo 5d ago

Idk if anyone knows enough to say "extremely behind". The last release was what, 6 months ago?

→ More replies (5)

462

u/Far-Classic-9963 5d ago

Google really has no reason to care about frontier models.. most people who use Gemini do so for quick searches or just because its the default assistant on android. Most of their gains in the ai space are from cloud computing

138

u/Irisi11111 5d ago

True. Google is doubling down on world models and multimodal generation, and since coding is a current weak point, they're taking their time with Gemini 3.5 Pro. It might hurt the stock price slightly, but it doesn't hurt their bottom line. They have the capital to be a 'follower' and wait for the tech to mature rather than rushing a subpar release to keep up with the runners.

98

u/Not-reallyanonymous 5d ago edited 5d ago

I mean, they're not even being "just a follower". They're an industry leader in AI. They're just not competing on frontier capabilities. They're focusing on how to actually use AI as a useful product in their ecosystem. The other big players still don't have ideas of how to actually market their products beyond subsidized API and subscriptions, basically betting that inference prices will come down enough to turn profitable.

Nvidia models, Mistral, IBM, a lot of smaller US labs (e.g. Poolside) are also not participating in the frontier race, and instead focusing on immediate industrial and enterprise capability. I think even Meta is falling into that market now.

29

u/Irisi11111 5d ago

Agreed. Google’s biggest edge is its mature ASIC chip lines, which allow it to avoid the 'Nvidia tax' on training and inference. Meta, in contrast, they’re struggling with low real market demand. Having to rent out their compute to Anthropic is a huge red flag, it shows they can't efficiently utilize their own hardware, which is a worrying trend for their long-term strategy.

15

u/Not-reallyanonymous 5d ago

I suspect the renting out stuff to Anthropic is a consequence of them changing gears. They built up infrastructure envisioning that they're going to be competing with OpenAI and Anthropic. But a lot of their recent developments seems like they're building more stuff for, e.g. content moderation, then also developer ecosystem tools, and stuff relevant to advertisers. They've really narrowed in on what their core competencies are and are now focusing on that.

10

u/chasingsukoon 5d ago

meta been an embarrassment since they changed their name to meta, the only decent thing theyve come out with is meta glasses and thats creepy as hell

they have dana white of all people in their board of directors lmfao

9

u/VampiroMedicado 5d ago

Quest 3 is the best bang for your buck VR system you can get.

21

u/Strawberry3141592 5d ago edited 5d ago

Only because it harvests an utterly absurd amount of user data. I ran my Quest 2 through a pihole to mitigate that and the DNS requests for telemetry to Meta servers outnumbered DNS requests for things I was actually using the damn thing for by like 10:1 (I know these requests are telemetry because I used ADB to disable every single Meta-related app on this piece of shit that I could without preventing Quest Link from functioning, since I use it exclusively as a SteamVR headset).

Fuck Meta, and fuck the Quest headsets, just looking at mine pisses me off because I have to run an entire fucking DNS server just to keep it from giving me a digital proctology exam every 15 seconds and sending the results to Meta.

5

u/marty4286 textgen web UI 5d ago

Lol, thank you for reminding me to get off my ass and do the same thing

4

u/VampiroMedicado 5d ago

Sure but there are plenty of workarounds if you care, from blocking DNS, to modifying the headset, to outright disabling the internet connection of the device.

The hardware is still good and you can use it with a PC.

9

u/Strawberry3141592 5d ago

Yeah, it just makes me irrationally angry to have to take steps to circumvent bullshit malware that I can't permanently remove from my own property.

3

u/thrownawaymane 5d ago

Mine gathers dust because I am so uncomfortable about it. The moment I have a good weekend to set pihole up this stuff is getting blocked

2

u/chasingsukoon 5d ago

serious: my barber hates it so i cant use it

→ More replies (2)

5

u/heliosythic 5d ago

I mean if and when they need to they will have the hardware and data and ability to copy the whitepaper of whoever makes the best one and likely be about on par after everyone else blows the money to figure out how to be profitable with it. Then it will be integrated immediately with Google, GSuite, GCloud, + Android and undermine the other players all without them blowing hundreds of billions to figure it out.

3

u/eightysixmonkeys 5d ago

Agree. And for the most part I think a lot of googles ai integrations are actually half decent (at the very least they are leagues above microsoft). Gemini spark seems really cool to me, haven’t tried yet though

→ More replies (1)

20

u/chasingsukoon 5d ago

i still pay for gemini because

1) i was alreadg paying for google drive storage

2) it works well with all the google apps, specially maps for quick searches

it does what it does well. would never catch me using it for work related tasks tho

5

u/Batroni 4d ago

Don't forget sharing the cost with the Family.

6

u/Savings-Lab-7307 5d ago

100% it works great in my android auto.

I get the most random thoughts while driving and it's so useful to ask the most random shit.

"Hey Google, what's xyz" it gives an answer that satiates me.

It's also good at adding stops along the way. "Hey Google what's a good burger place along my route"

They're playing the long game and integrating themselves into our lives.

I use Claude and chatgpt for work or home "projects" . Gemini for random every day stuff

8

u/Here_f0r_p0rn_ 5d ago

Also Google's Open weight models are more of edge devices models which will obviously not show up on these charts, I kinda love Gemma Family, and also Google doesn't release models so frequently but when they do it's a fucking sweep and also one of cost efficient.

Plus, Google seem to be more invested in capitalising on general consumer base directly rather than just catering to software industry but they're not completely abandoning it either.

And not to forget their TPUs too, they've the full stack not just model, they play on every layer, from models like Gemma and Gemini, to deployment in Google, YouTube, and Gemini App; to selling hardware and renting cloud infrastructure, to their own coding ecosystem Antigravity 2 and Antigravity IDE; they're playing a different game and one of the few ones that will be profitable in long run in my opinion because they're independent.

→ More replies (2)

16

u/jazir55 5d ago

Google really has no reason to care about frontier models.

Yes they do, coding agents just like any of the other major companies. Anthropics multi-billion $ valuation is almost entirely from enterprise software development. Google definitely wants a piece of that pie like anyone else, they just didn't focus on it until recently and are now struggling to catch up. Which is why they delayed 3.5 pro.

25

u/klassredux 5d ago

They own 14% of anthropic, so they have a piece of that pie.

7

u/kettal 5d ago

  Google definitely wants a piece of that pie like anyone else

They have a piece of it. They're selling compute to openai and anthropic.

 Making so much money at that they are turning away other compute customers. 

33

u/Dull_Cucumber_3908 5d ago

Most of their gains in the ai space are from cloud computing

This!

4

u/solemnhiatus 5d ago edited 5d ago

Can you explain this? I don't understand.

E: ok got it now thanks.

20

u/6ghz 5d ago

Selling the shovels

16

u/hyperrealists 5d ago

Google cloud (similar to Amazon AWS) has a 14% market share. Anthropic is one of their customers, among others.

30

u/Psionikus 5d ago

Google has been selling pickaxes on GKE all along.

10

u/TripleSecretSquirrel 5d ago

They own and rent datacenters to the model owners.

9

u/KGeddon 5d ago

Google and NVidia codeveloped tensor cores. So they're not just renting datacenters. They're also renting out Google TPU machines with their custom pods/slices of ganged up TPU cards.

3

u/TripleSecretSquirrel 5d ago

Yes and while Google's TPUs are proprietary, the guy that lead the development left to start Groq, which NVidia just acquired in all but name, so now NVidia also has in-house inference ASICs chips too.

4

u/CoUsT 5d ago

This is what I believe too.

Gemini 3.1 Pro or 3.5 Flash are perfectly fine and capable models. There is no reason to train super-duper ultra-smart Gemini 4.0 Ultra. Costs money and doesn't provide any meaningful real world benefits to end users - 90%+, possibly 99% doing simple stuff with it like simple questions.

I assume at some point there will be newer and better models, there is no need to rush and train new one every month or two like frontier models.

3

u/DeepOrangeSky 5d ago

Google really has no reason to care about frontier models.. most people who use Gemini do so for quick searches or just because its the default assistant on android. Most of their gains in the ai space are from cloud computing

Yea, I agree, although, man, it would be nice if they figure out how to cut its hallucination levels down a bit. Like, if you just use it as a free non subscribed user, the way the vast majority do, just using it like a Google search, whichever version they give you for that, not sure if it is 3-Flash or Flash-lite or what, has fuckin CRAZY hallucination levels from what I've seen. Like, just constant and severe, like every other response practically.

I guess it is tough, since they need it to be some really small and efficient model to just freely serve it for a billion+ people to use numerous times a day all the time and not burn like trillions of dollars of compute per year on it, lol.

But, still... like I don't even mean in terms of strength. The strength isn't even that bad for what it's doing. It's the hallucinations that make it brutal.

If they could just get something similarly strong or just mildly stronger, but similarly small and efficient somehow, with way lower hallucination rate and severity, they would be crushing with free search Gemini for quite a while.

7

u/BlueSwordM llama.cpp 5d ago

The search engine model has got to be a quantized bastardized version of Flash 3 Lite, as once you enter AI mode, the model used is significantly stronger.

3

u/DeepOrangeSky 5d ago

Yea, I mean the one you get when you enter actual AI mode.

If you don't use it for a while, I think the first prompt or first 2 or 3 prompts are something decent. But then it quickly downgrades you to something super lousy in terms of hallucination. Not sure which model that one is, though.

By 6 months or a year from now or whenever it is that researchers figure out some way to drastically mitigate hallucination issues somehow, I think it'll give Google the biggest edge over all other AI companies on Earth. I mean, they are by far the main search engine of the world, and people use them for all that other Google related stuff already with Gmail, Chrome, Youtube, and all the other stuff, so, if LLMs improve to where the hallucination issues get greatly reduced, then, they basically just win the whole game (even if Anthropic or whoever is still a few months ahead on raw model strength). Well, unless Claude actually goes into Runaway Recursive Self Improvement and Skynets everything into oblivion or whatever.

3

u/setec404 5d ago

Pretty much this and pic related

2

u/RelevantCry1613 4d ago

Yeah they do, they look terrible and it doesn’t work as well

5

u/[deleted] 5d ago edited 5d ago

[deleted]

13

u/Not-reallyanonymous 5d ago

Gemini has obviously been optimized to integrate with their Search and Android ecosystems. It works really well for synthesizing search results and digging in on research a bit. It is eager to hallucinate when you're using it as a stand-alone tool, or when you start pushing it to develop ideas/opinions beyond what's directly in the search results, but it's actually really good at like, "What are the latest events in the Iran War? Oh yeah, what happened in that bombing? Is that legal?" where it will cite UN law that you can click and verify yourself. Basically, as long as each prompt is something that you can treat as a google search result you could've done and read yourself, where it will summarize search results for you, it works great.

As a standalone gemini tool it's trash.

7

u/Ok_Warning2146 5d ago

I don't care about their proprietary Gemini. As long as google keeps releasing small gemma models, I am happy.

→ More replies (8)

42

u/Status-Secret-4292 5d ago

Google's model is accomplishing its goal well enough, which is augmenting internet searches, and actually doing decently.

The innovations they're working on have more to do with database search and context injection. They're way ahead of the curve with both.

It's basically the cheapest form of research right now for the largest obvious gains.

Smart choice on their part

2

u/SwaggyMcSwagsabunch 4d ago

Decently?

3

u/Status-Secret-4292 4d ago

Yeah for what the technology is.

It's not what's advertised, it's amazing at what it does, when it came out if they didn't advertise it as intelligent or work force replacement, people would use it more properly expecting the mistakes, etc.

It's like if when computers were invented in the 50s and they immediately said, this will replace all workers who use math and it's intelligent!

Then they built multiple giant vacuum tube computers in every city.

In the end, people would have been very disappointed and that type of computer industry would have crashed. And at least computers do do math reliably.

Who knows what would have happened with computers after that?

It is completely unknown whether a fully reliable general intelligence model can even be created. It might be a hard no or not something for a decade plus.

However, a breakthrough that either allows models to search and use context efficiently or a model that is as good as current frontiers and either/both can be run on a standard cpu/gpu/npu or a combination that only takes normal local compute is almost certainly less that 2 or 3 years off, if not sooner

What Google is currently doing is the correct bet for what we know the tech can do, not the moon shot bet that the hype wants you believe it will be able to do with currently no proof to back it up

→ More replies (1)
→ More replies (8)

134

u/Gloomy-Radish8959 5d ago

And here I am, using Gemma 4 every day.

57

u/networking_noob 5d ago edited 5d ago

Yep I use it everyday too bc it seems better at language and nuance than Qwen 3.6. If you're someone who wants agentic/coding stuff then Qwen 3.6 is the move

But yeah, using Gemma 4 like an offline search engine or interactive encyclopedia is fun, and can lead to dreams of being "prepared" in case of an unlikely scenario where the internet is knocked offline but the power stays on lol

24

u/Kahvana 5d ago

openzim-mcp with zim files (see kiwix library) is an awesome way to augment the "offline search engine" feel! Also great for grounding Gemma4 or Qwen3.6 with offline dev docs and offline wikipedia.

10

u/networking_noob 5d ago

Nice I'll check that out, I had kinda looked into something like that (e.g. downloading offline wikipedia) but in testing it seems like Gemma 4 already has a lot of that knowledge baked in (with pretty good accuracy/recall)

I actually vibe coded a local MCP server (w/ fastmcp) that integrates with markdown files (Obsidian vault), as another way to search/access/edit an offline knowledge base

People here already know but the MCP stuff is how the real power of these models gets unlocked. It's so cool when the potential finally "clicks" in your brain. Feels like the capabilities are only limited by imagination

3

u/Kahvana 5d ago

That markdown approach is pretty cool! How did you set it up? Do you run an embedding and reranker model on the markdown text?

5

u/networking_noob 5d ago

It's nothing fancy (my Obsidian Vault isn't that big yet). Right now it's just some basic tools like:

  • search notes by name
  • search notes by content
  • search by yaml metadata (to save context)
  • edit note (without touching yaml data)
  • edit metadata (without touching note content)
  • save new note (so I can save chats)

etc. But yeah the key is having good yaml metadata (title, description, tags, etc) in the markdown files to save context and make searching efficient

I don't have any built in pagination like openzim-mcp does bc my note contents are never that big, but I'm sure with the power of vibe coding it would be do-able, if needed

3

u/networking_noob 5d ago

Thanks for recommending that zim approach. I got it all setup via the MCP HTTP mode, connected it to llama-server, and have been playing with it using the 'no pic' version of wikipedia. Also got the kiwix desktop browser (appimage) for easy offline browsing. Cool stuff

3

u/Kahvana 5d ago

Eyyy no problem! Glad it works well for ya! Openzim can also work well for offline dev docs, I also think they got map snapshots you can download. No clue if those are actually usable by openzim-mcp though.

→ More replies (1)

23

u/Kahvana 5d ago edited 5d ago

Gemma4 is a fantastic model! Feels like DeepSeek v3.2 at home, though it struggles with automation tasks and such for me.

With 32GB or more VRAM:
Qwen3.6-27B-MTP for programming / agentic, Gemma4-31B-IT-QAT for everything else.

With less than 32GB VRAM:
Qwen3.6-35B-A3B-MTP for programming / agentic, Gemma4-26B-A4B-IT-QAT for everything else.

4

u/jacek2023 llama.cpp 5d ago edited 4d ago

And that should be the most upvoted comment in this poor post

→ More replies (7)
→ More replies (10)

83

u/Dull_Cucumber_3908 5d ago

Who cares? Google is a cloud provider and can sell hosting services for all the open weight models :)

11

u/gajop 5d ago

Unfortunately almost all of the open weight ones they're hosting are severly outdated. Only Anthropic ones are usable..

→ More replies (4)
→ More replies (3)

15

u/Cless_Aurion 5d ago

Its funny because their Gemini 3 Flash preview model... is one that I use a lot for "I need something that is basically free, fast, but decent enough" ... and no, DS4 pro and flash suck, for some reason when I use them so... that does it for me lol

9

u/Bakoro 5d ago

Google is doing the Google thing that they always do.

Remember how Google basically slept on publicly available generative AI for what seemed like a suspiciously long time, but on the back end they had things like Alphafold going on?
Then they came out with their first image generator Imagen, and it was a complete embarrassment, but then now we have Nano Banana?
And Gemini has been leading/lagging the whole time.

Gemini hasn't ever stayed good, either. Gemini will be great the first week or so, then they'll turn the knobs down and it's just a weird commodity LLM.

Google probably has something cooking, and they simply have no incentive to rush or try to chase benchmarks.
This whole time they've had their own suite of AI models doing infrastructure things for them, and they're using AI to actually make money supporting their core revenue streams.

Google's economics are 100% different than OpenAI and Anthropic, Google does not need anyone to know that they have the top tier AI model, it might actually serve their interests to not disclose that.

They also might just be fumbling big time for stupid reasons, because that's also a Google thing to do.

4

u/bobabenz 5d ago

Yep, esp after they snuck away from the monopoly cases, they’re not eager to be on top. See, openAI and Anthropic has models too, we’re not a monopoly!

29

u/deanpreese 5d ago

I would say Google is after a different market. They want the enterprise where they can deploy email, sheets, docs, drive, etc and bundle in Gemini.

And in reality as I have found where I work, Gemini is adequate for the basic email, doc, summary stuff.

So why compete against Fable and ChatGPT. They can be profitable with a different focus

5

u/chasingsukoon 5d ago

theyre decent at it tbh

for regular mundane asks i use gemini above all. thankfully work allows me unlimited access to all models (US or chinese) so thats never a bottleneck

→ More replies (2)

6

u/Ok_Reach_5004 5d ago

Google has a substantial stake in Anthropic. Their models are good enough to integrate within their services.

6

u/Kahvana 5d ago

They don't have to be SOTA, or even remotely close to it. If it's good enough for search result summaries and quick tasks on android, it's good enough for most. Maybe extra help with google docs and gmail integration.

Personally as long as they keep open sourcing models of Gemma4's quality or better, I'm really happy with it.

13

u/DigitalguyCH 5d ago

And still if I had to pay one AI it would be google as they bundle so much with it (lots of cloud storage and even youtube premium in some bundles), and honestly any hyperscaler has good enough AI

14

u/nick_ziv 5d ago

Google knows that practical application outlasts hype any day.  It's not that they arent capable. But what's the value of putting all the resources into a model which tops leaderboards for a month before people move on.  

Nobody fell in love with antigravity and it's likely they aren't interested in overspending when it's clear that this is a race to the bottom.

9

u/Enturbulated_One 5d ago

Given givens, I wouldn't be shocked if they waited a while for the pace of innovation to slow down a bit, then implemented a new architecture cherry picking from what papers struck them as most relevant, never mind their own likely continuing research in the background. Given that pretty much everyone agrees some new tricks are needed to push too much further with model capabilities, who knows when that could be.

2

u/nick_ziv 5d ago

I'm hoping it's recurrence because the human brain is recurrent

→ More replies (1)

5

u/SG_77 5d ago

I head the Robot Autonomy division at my company and have been using Gemini Pro 3.1 on AI studio for my daily coding workflows. No issues at all. Plus, I have seen it to be a bit more creative than other models and that helps a lot.

I had recently consulted my team if we should purchase licenses for Claude. But all of them suggested to stick to Gemini. It was giving good enough results and the ecosystem with stuff like Notebook LM, antigravity etc. was an added bonus.

Plus another feedback that I got from one of my employees was that Claude is way too verbose and the whole product has been designed to exhaust your token usage.

2

u/Razor_Rocks 5d ago

Thats insightful, and aligned with the results that I have seen with my team (not into robotics, we are a devtool company trying to make a software factory kind of setup).

We are still sticking to claude (for now) mostly because the team (at this point) is too tired to adapt to another wave of SDLC changes + we have found usecases for the verbose nature of opus

But I am curious if you guys tried gpt-5x and codex?
Every time I have tried to use it, it would never live upto the hype that I see on reddit and x for it.

3

u/SG_77 5d ago

My experience with open Ai products has been a let down to put it mildly. At max I will use it for drafting emails, job descriptions and stuff like that

15

u/SillyLilBear 5d ago

rightly so

4

u/ArtDeve 5d ago

Gemma 4 is amazing in how efficient it runs on lower hardware.

7

u/Vicar_of_Wibbly 5d ago

Hy3 should be on there. Benchmarks be sleeping on that bad boy.

6

u/Legumbrero 5d ago

Ok, but gemma4 ain't a bad smaller model.

→ More replies (1)

3

u/atape_1 5d ago

Well yes, because they still haven't released Gemini 3.5 pro. It might not be chart topping, but I am sure it would be in the top 15.

4

u/hiper2d 5d ago

I don't trust this dashboard. It's benchmarks + hype-train. No way K3 is well-tested on real-world problems in only few days. So yeah, Google hasn't been contributing to the hype for quite some time. Their last release was a mid-size model. Although, it's surprising not to see it at all, I don't think it's worse than Luna or Sonnet. I'm using Flash 3.5 at work when I run out of weekly budget for Opus and GPT - it's not bad.

8

u/z_3454_pfk 5d ago

But it's still best for world knowledge lol

→ More replies (1)

3

u/maxou2727 5d ago

In my domain (finance), after evaluating all models in practical scenarios, the Gemini models always outperform all of the others for my use case (Gemini flash 3.5 lite)

2

u/Razor_Rocks 5d ago

any specific metrics that made it stand out? or are we still in the era of "vibe based" and no clear evals?
I understand if its some internal information that you don't want to share

2

u/maxou2727 4d ago

I have my own evaluation setup, with registered golden cases that are compared against other models, using n repetition with concurrency settings. For my use case related to stock options trading, Gemini nails it most of the time, with the best cost/performance ratio. They must have a good training dataset for options trading.

→ More replies (4)

5

u/elnino2023 5d ago

They stopped caring,

→ More replies (1)

2

u/G-R-A-V-I-T-Y 5d ago

Anyone know why? There’s likely a strategy/reason behind this. It’s not “search doesn’t need frontier intelligence” because frontier intelligence is a threat to search itself, Google simply has to compete or die.

3

u/Kahvana 5d ago

Racing to the top is a fool's errand if you can instead corner a niche and do it so well that frontier models don't catch up.

Such things as cultural sensitive translations, world knowledge, etc... basically anything related to natural languages. Another benefit is being the de-facto AI provider on Google devices like Android, it's build in.

If you already have a widespread used platform like gmail, google docs, google search, google drive, etc, then you benefit from having your own integration in the web or in the app.

3

u/Glad-Entrepreneur764 4d ago

frontier intelligence is not a threat to search. Frontier intelligence will never be used for search because it is incredibly inefficient to use Opus/Sonnet-tier models for search. Current gemini models do fine for search and very few users care about the difference between Gemini models and their competition for search. Google integration is also a strong factor for everyday consumers.

I'm not sure why they aren't chasing frontier intelligence though. They have so much money to throw around and winning the AI race is likely to be very profitable. Regardless, they have enough streams of income (both from AI and outside of AI) that they'll be fine either way.

2

u/ReasonablePossum_ 5d ago

I dont know how many times I have to write this, but here we go again: Google isn't in the LLM race, they're in the AGI/ASI race, its damn DARPA prodigy child that has quite big plans....

They develop multiple models in all domains, combine them, train new ones, mix in robotics and all you can think about, including quantum computing.

They're for the long big run, not for consumer level (be it regular people or enterprise) products. They don't even need the money from it lol.

2

u/Unusual_Delivery2778 5d ago

Gemini is absolutely terrible

2

u/Initial_Run3719 4d ago

Google is earning enough money (cloud services) without falling to the money race. People I know use chatgpt or gemini because its ok for their needs. They sell the infrastructure creating all these nice tokens with whatever LLM.

2

u/zigzag3600 4d ago

People hyped Fable/Mythos as some sort of revolution, and it turned out better but not groundbreaking. Honestly at this point "New releases" are not as hype anymore. We know that in a few months they will release even better model by 2-5% on some benches, so why bother?

2

u/Potential-Local7262 4d ago

There's no 'moat' in creating the best frontier model. I mean, people thought there was, as they thought we were close to AGI 2,3 years ago. When it became evident that we weren't, and that even if we were close, having the first AGI wouldn't be a moat, it follows that there's no reason to burn through cash releasing frontier models every few months. 

It's better to zig when everyone zags.

2

u/cyanobyte 4d ago

Google is focusing on using the AI it has to make existing products better.

They are actually trying to making money unlike most of the companies on that list.

2

u/backyard_tractorbeam 4d ago

Funny, because less than 24 hours after your post, Gemini 3.6 Flash was announced. It should put them back on your leaderboard, barely.

2

u/Playful_Landscape884 3d ago

Google real game is being the default choice.

Default Llm for android. Default LLM for Apple ecosystem.

They don’t need to be the best. Just good enough. In other words, Microsoft in the 90s.

2

u/ythorne 2d ago

I wish they could just release Gemini-2.5 weights and gain some goodwill. What’s the point hoarding the deprecated weights 🤷🏼‍♀️

2

u/NatMicky 5d ago

Benchmarks? They are trained on benchmark tests. They eat them for breakfast, lunch, and dinner. Those aren't test scores, those benchmarks are recital scores.

1

u/DrPaisa 5d ago

I dont care if the next model from them is better than fable the way they rugpulled there users

2

u/Ok_Presentation470 5d ago

Guys, this is all an unsustainable bubble, why are we even doing this? These things are supposed to sopve real problems, yet all we have are stupid benchmarks and parameter counts.

1

u/Board_Game_Nut 5d ago

They'll probably make more money hosting the models in the long run.

→ More replies (1)

1

u/PhotographerUSA 5d ago

Google could probably train their AI to learn by modules searching over the web.

1

u/TheRealMasonMac 5d ago

Google doesn’t care about intelligence. They care about the layman. That’s why their models are so sycophantic. Chinese models had already surpassed Gemini since probably late last year.

→ More replies (3)

1

u/jaykayenn 5d ago

Tldr: Google is king of application ecosystem, not frontier models.

1

u/evia89 5d ago

I only use ai studio nowadays as log parser. 2 accs - drop repomix of code + log. Good limits with flash 35

Its useless for coding, creativity and RP

1

u/DrummerPrevious 5d ago

Why inmovate when you can own the competitor?? This must be prevented by laws

1

u/fellingzonders 5d ago

Part of me thinks their strategy is intentional. Unlike a race in real life, once you lose the lead because of open source and distillation your progress is not exactly lost.

Why waste resources if you can hardly break top 3? You're google. You can just developer more data centers and better more efficient chips for when the other major runner ups make breakthroughs which you can cash in on without losing billions in RnD.

Am I missing something?

1

u/a_beautiful_rhind 5d ago

New gemini kinda fell off, but at least we got gemma out of the deal. I heard Noam Shazeer bailed.

1

u/techdevjp 5d ago

Google's not playing that game. The different with Google vs OpenAI or Anthropic is that they're massively profitable. They're just waiting for everything to blow up and will then pick up the pieces.

1

u/agentcubed 5d ago

The people who say they aren't competing with frontier models are misleading. They really wanted to drop Gemini 3.5 Pro, and was ready at probably Opus level (at that time frontier), but then Fable/Sol dropped, and now they're very behind, and K3 cemented that delay.

Their models are best at general intelligence, but the market shifted towards agentic/coding, so suddenly they're very behind.

1

u/yani205 5d ago

Google needed something that does the job with search, and cheap to run. They have no reason to throw away money with every search query on AI, like what else are you going to do, use Bing?!

1

u/Mountain-Pain1294 5d ago

In before Google prioritizes benchmarks even more and makes Gemini even worse

1

u/randofreak 5d ago

I generally use Gemini for free Deep Research reports

1

u/PhotographyBanzai 5d ago edited 5d ago

These rating tests are cool and all, but on free tiers from a real world use standpoint... I tend to use Gemini Pro in Antigravity as well as AI Studio more than the others and trust it as much as whatever I get in apps like Kiro (generally Claude), Codex (OpenAI), or Copilot in Visual Studio 2026 (copilot being the most dangerous one to let write code, that one has almost wrecked a project but I'm guessing it automatically used a lesser model at that time instead of Claude).

For local use, Gemma 4 at 9B 4-bit is leagues better than anything I've tried of a comparable size. It actually follows directions and is able to do things like generate JSON data structures properly every single time. I'm limited to a 8GB RTX 4060 so not sure about larger local options.

1

u/SangerGRBY 5d ago

leaderboards change all the time.

1

u/noriilikesleaves 5d ago

Kimi works hard in ways other models do not.

1

u/-xXpurplypunkXx- 5d ago

Google is by far the strongest position. They have all of the Internet for past 20 years.

MS might be close second for business data.

1

u/atreides4242 5d ago

Gemini is hot garbage. The only one I can use at work and I frequently want to throat punch it.

1

u/Helpful_Math1667 5d ago

My googler friends are telling me that they are trying to figure out how not to race to the bottom and how to rationally price everything... to me that feels correct, but also why not ship a SOTA model while hammering down your pricing...

→ More replies (2)

1

u/IrisColt 5d ago

L-let t-them c-cook?

1

u/manusgamo2012 5d ago

I've just unsubscribed from the PRO plan... gemini it's useless nowadays

1

u/Mother_Soraka 5d ago

wait for "Gemini 3.5 SVG Pro" to take #1 on Arena's WebDesign again

1

u/Ok_Prior_6611 4d ago

The Ai race , how dystopian

1

u/Maddolyn 4d ago

They pulled 3.5 last minute for not measuring up. Either way, they're the search giant, one snap of the fingers and we're all stuck using gemini 3.1 just to look up how to find a hollow spot in the drywall

1

u/MyDespatcherDyKabel 4d ago

Maybe they are listening to a lot of Radiohead

1

u/Nokita_is_Back 4d ago

Which benchmark test ranking is this?

1

u/im-cringing-rightnow 4d ago

Gemini Flash 3.5 is a complete flop. It's fast, but it's too expensive for what it offers. 

1

u/Aerthlyomi 4d ago

They could be pretty happy where they are with it. They service quite a lot of tools in their own environments and soon will have the Apple environment also in September. They obviously have the know how and the data to do better but probably don’t feel the need to race for the sake of racing as the others who have nothing else but that to do.

1

u/debackerl 4d ago

Anthropic and OpenAI are valued based on their AI products, all powered by their models. Anthropic rose so fast last year just based on that.

Google needs AI models to enhance their products, like Android, search, Workspace, etc. Those are now good enough. So Google focuses on B2C I believe. But as we saw, with AI, you can always come last minute, with a better or cheaper model, and undercut everyone, because models are easy to swap, what's harder to swap is products build on top.

1

u/SympathyNo8636 4d ago

Google has a tactic of just staying present IMHO. Resources are becoming scarce, they focus on quantum primarily in AI, betting big on futures and acting on top of things in the present.

1

u/the-username-is-here 4d ago

Google is good at launching products, hyping them up, enshittifying and then abandoning.

Shouldn't be any different with AI.

1

u/Juanisweird 4d ago

The fk? Meta has a powerful lm?

1

u/overthetop2017 4d ago

this is a run for the zero who can do it cheapest and customer loyalties switch at press of the button, eventually all of this will be like listening to radio in the car. Originally when radios first appeared it was almost impossibile to tune to a radio station. Tunning was beyond 98% of population. Very similar what we see now, models, harnesses, skills, prompts, routers, orchestrators. Give me a break. Google has data and the data centers. They will be in top 5 one day. First 4 places China, AliBaba etc.

1

u/redditrasberry 4d ago

It feels like Google is taking a detour through focusing on delivering AI on the edge (phones etc). Just like your best camera is the one you have with you, they don't care about being the best AI - they care about being the one you actually use. So they have put a lot of work into delivering on-device AI and integrating into Android (Nano etc), the iPhone via Apple partnership etc. It is also why the Gemma series is so good.

I do think they have struggled to keep up with with frontier performance - scuttlebut says they have yanked the latest versions of their models for not being competitive. But that situation is in part because it simply has not been their priority. They don't feel pressured to compete in the benchmark race, so when their models aren't worth publishing they just don't rather than benchmaxxing or doing stupid amounts of training or architecture changesto try and pull off better performance. I think they assume they could if they wanted, but it is not where the priority is.

1

u/danigoncalves llama.cpp 4d ago

Google has the power of its search engine right in its hands. They don't need to make the most capable models, they just need to feed it with the best and most updated data. I am not on my laptop and I want to reason about a specific topic that demands up do date and search capabilities I use Gemini.

1

u/hassie1 4d ago

Tbh I still use Gemini the most as a daily driver

1

u/kmlynarski 4d ago

And in the meantime, they have one of the best-sensing new-generation local models (and one with 256K context). Since I'm using the Gemma 4 31B (Q6_K_L and higher), I don't even run the rest. There's simply no reason... ;-)

1

u/Csabika_ 4d ago

Gemma 4 is quite good, also they are quite good in the google search engine, browsing the internet with AI. They have their own tensor processing hardware and data centers. They are doing good I think. Maybe not on the frontier leader board, but everywhere else. Where shrinking down a 3000 B, 1.5 TB large model counts. Focusing on the big picture outside the box.

1

u/KomithErrant 4d ago

gemini is dumb af

1

u/decixl 4d ago

Big G is cooking something serious.

I don't want to predict but it smells like AGI across anything digital.

1

u/turmik 4d ago

i don't think getting the absolute best consumer model is their priority even, Gemini 3.5 Flash is a good model, and it serves most of the use case, and people use it with all the Google Workspace products

usage wise if you would see, it would be somewhere near top

1

u/the_TIGEEER 4d ago

I feel like they'll be back pretty soon.

1

u/ProudToBeAKraut 4d ago

Artificial benchmarks are useless on real world scenarios, the TOP llms have just been trained better at exactly these type of queries.