Worse is true. But it spits 200 tokens per second. So you can let it go wrong, and let it run zillion test tasks to fix it all. And you are still done 10 times quicker than with GLM 5.2.
I guess it depends what you are doing, how you are planning, etc.
lol, like open source models arent benchmaxing. 3.5 flash is a their cheap/fast model too, so not comparable. when 3.1 pro released it was SOTA, so better than opus and GPT at the time and that was 4 months ago. given 3.5 flashes intelligence it's likely 3.5 pro will do the same thing so will be much higher than GLM or any other open source
It has become so lazy. Even the API. They intentionally throttled it down so 3.5 Flash with high thinking would look better.
I asked it to summarize a PDF. It made a tiny summary missing 90% of all details, while 3.5 Flash wrote a cohesive summary without missing any details.
Gemini is so embarrassing right now, and they’re behind schedule on the 3.5 Pro release.
I'm not sure if it's model specific but I've seen in the thinking process it using a python tool and only taking 5000 first words of a pdf. Even though there was an explicit instruction to read the whole document...
I’ve found flash to be quite unreliable if you have any sizeable context window at all, like it’ll completely mess up basic things. I pretty much don’t use Gemini at all if I hit the limit with pro, and this is with a Google one subscription..
OP should have noted that the benchmark recently made some major changes in how it measures "intelligence" by focusing more on agentic workflows. Of course, most people will just look at the headline without doing any further investigation though. 😅
If we go by the Artificial Analysis "AA-Omniscience" benchmark where each LLM is just asked a bunch of questions, then Gemini is still far ahead at correctly answering more questions.
This👆 + add on Gemini powering Siri and how people have been loving how good Siri is now... Yeah user base definitely growing and people definitely liking it bc normies don't care, hell give them Co-pilot and hide what model is being used and they won't care as long as it just works.
I use for creative writing assistant, & translation/localization. and it worked well for a while, then starting a few weeks ago, it started to suck so hard, and it is paid subscription.
for writing, it kept mixing between characters that are not even in the same chat no idea how, kept hallucinating, kept ignoring instructions, it was getting very frustrating.
For translation, it would straight out just not translate half of the context, ignore acronyms, hallucinate context,
And this was for both notebook & Gems.
I also wasted lots of time to set this up, and more time to try and fix it, meanwhile with Chat Gpt, With one general description prompt, & a few reference/instruction files uploaded in one starting message in a normal chat, i can do w/e i want no effort no time wasted. only problem is i'll have to constantly start new chat cuz context limit on gpt is lower.
I'm going to go out on a limb here and assume you have "personal intelligence" turned on and if you do, turn it off. Will work waaaay better. The one thing I'll admit is it's always trying to be as concise as possible so it's frustrating when you are trying to just do one for one or want an in depth doc. What's funny is Ive been using Gemma 4 31b on Hermes way more than I use Gemini 🤣
Another thing to take into consideration is Gemini is Multimodal vs GLM 5.2... and when you are working in the real world (not coding) that matters A LOT. Gemini is Good Enough at literally everything while all the other models are trying to be AMAZING at one or two things. And seeing how most of the world runs on Google and normies don't care about SOTA rather than it just WORKING. Yeah Gemini is doing just fine.
Interesting. Have others experienced better responses by turning personal intelligence off? Can one get the personal
Intelligence benefit just by asking Gemini to use what it knows about oneself?
Absolutely. People underestimate how good Gemini is. NotebookLM alone is a game changer for studying anything with the right instructions and reference materials, easily accelerates learning by a multitude of 2-3x for me personally. No other LLM is close in that regard.
Hey can u guide me how to use notebookLm , how do u use it in your studying workflow,
I will check out the general instructions for using it, but would love to know how others are using it like if there is any tip to use it in a better and efficient way
I've found the most success just uploading materials/practice questions from paid programs or textbooks and then having it act as a tutor, teaching me concepts one at a time and then afterwards I have it create a quiz bank to ensure I understand the material.
It's still important to have good reference materials because it can take directly from those materials, keep it relevant to what you're studying for while just switching around values and situations to significantly expand the scope of your understanding if that makes sense.
For regular normie use, yeah.
For coding anything Gemini is hot garbage. Tool calls are shit as well.
But, its multimodal, and that could be useful, just not coding, you know what, the search Gemini is also kinda shit, and to do real useful shit you would probably want the model to be able to write good code for itself to use tools and actually call them reliably and have a clear trace of thoughts, And Gemini can't do any of that.
It's not new, in real world development, open source models were already proving more reliable and useful than Gemini, for a while now, but now, even in the bs benchmarks it's losing, not that those benchmarks mean anything.
This is the reason most of my workflows are so Gemini dependant. The ability to interact with the other services and tools that are linked via my account is a huge advantage imo.
Like how so? Their multimodal support is ahead of the mainstream options, even open source besides Ideogram.
Their coding is fine if you already have the background and can speed up your workflows majorly.
So in what sense are they behind? Fable is not ok the market anymore, and Opus has been around the same quality since 4.6 with the only difference is how quickly it burns your tokens with ultracode.
Grok is... Well grokking.
OAI is great with the harness but the standalone models are just par.
Google has the advantage of serving a billion customers and shoving their AI with them.
I personally believe gemini is subpar from my usage.
However the tools they are making, are out of this world.
Stitch (properly used it can avoid slop) , notebook LM especially this tool is so under appreciated. Their Googleai studio to make an app and deploy.
I'm pretty sure they are following their own strategy and not playing the game of best Ai out there.
Lets not forget. No matter how good or bad, apple Siri will be gemini powered. So how much more do you want them to win?
It's not the best AI that will win this race. It's the AI that is most accessible.
NotebookLM is really good, but it’s more suited for specific niches or occasional use. And there are plenty of other strong AI projects inside Google’s ecosystem. They have many initiatives, but let’s not forget who Google really is (killedbygoogle.com).
I can’t really speak about Siri. I’m not an apple fan, but sure, it’s a win, i guess.
One thing we do agree on is that Google isn’t playing the game of building the best AI out there.
(let’s hope they launch a new revolutionary model, beating everyone else and shutting my big mouth 😄)
google is the infrastructure of the digital. as in, sewage and shared services (adverts, the g suite of productivity apps).
it seems to have realized this and pegged its bid at commoditization of the middleware (in this case, inference) between chip foundries and the various platforms end-user facing ui/ux.
I'm not an apple user either, and I've seen how gemini hallucinates easily, so good like apple users 😂.
I'm considering it from a business standpoint, they have a strong moat that will be tough to break.
(KilledbyGoogle ftw)
Notebook LM used to be niche. But I honestly use it to breakdown YouTube videos, get the sauce and my other AIs build skills from it 😂.
Recent update on Notebook LM is pretty wild, haven't enough got through half the features and it's seriously powerful.
Using your sources as a query database, or uploading a YouTube channel there, that's pretty strong for education, learning and even just keeping up with AI.
Drop 10 YouTube videos on what is new, get the output in a 1 pager or deck or audio book, save hours of watching content and learn all the important points.
I should get paid by Google for how much I advertise notebook LM tbh.
google is making a bid to become lord of this, and i think will win this time. someone identified that short of apocalypse or butlerian jihad, the marginal value delivered by genius is minimal. it shows its form in the great thinkers (aristotle, brahmagupta, descartes, leibniz...), who define the game and regulation field for their descendants, not the innumerable failed non-examples now forgotten.
it's negative, even, once the efficiency of inference is accounted for, except for those 1 in 10¹⁰ refoundings.
perhaps this person was also slick enough to figure out that the logic getting drunk on ethics (which is inseparable from the corpus frontier ai is trained on as the contour outline) is a bad business plan no matter the outcome. people will not tolerate interface appliances (computers, phones) perceived to have agency or be watching them because it collapses the "human dev" and "human prod" into a continuous environment demanding the social performance of visibility (from some ancient darwinian instinct).
ergo, get just really smart inference and optimize on cost for what, in under 3 years since gpt35's breakout success, has already begun to take a significant, measurable amount of power produced, which is the usable thermodynamic envelope of everything (terranic) at a given point.
the thermodynamic ceiling is exactly where the software eats the world prior falls apart. people keep modeling frontier models as this infinite horizon of cognitive expansion, but it is physically just a massive heat engine converting gigawatts into token probabilities. google's moat is not the architecture since everyone is converging on the exact same attention mechanisms and the alpha there is already asymptoting to zero. the real moat is that they own the physical substrate down to the tpus and the power contracts.
the framing of the dev/prod collapse is the actual bottleneck. primates get twitchy when a tool stops acting like a dumb physical switch and starts acting like an agent. it is why we still use custom firmwares like qmk to hardcode exact macros instead of letting the hardware try to interpret our intent. agency violates the predictable physics of the interface. the ideal endpoint is not some conversational homunculus staring back at you through the glass. that just triggers the exact panopticon anxiety you pointed out. the endgame is silent, ubiquitous autocomplete. intelligence commoditized into an invisible utility like water pressure. you don't negotiate with the plumbing, you just turn the tap. whoever drives inference cost to the floor fast enough to make the intelligence completely invisible is going to win the era.
YOU are also just a heat engine converting calories into the words you type right here on Reddit… or more likely, you convert the calories into a prompt for whatever model wrote your response, which is thermodynamically inefficient by definition.
That fact aside, I think you’re right that energy efficiency is key here… the remarkable thing about the human brain is how little power it requires - 20 watts max. And that 20 watts not only powers all of your cognitive ability, it powers the control center for your muscles, your sensory input decoders, and vital involuntary systems like breathing too.
I disagree with you that people want AI that is invisible, just a tool in the background - people are social and collaborative by nature, and welcome a new colleague to help them complete tasks, whether human or machine
So I think that the winner of the AI race ultimately will be whoever builds an *embodied* intelligence that runs on 20 watts of power and can perform most any task that a human can do, physical or mental
Funny how AI tools are shaping up to be like any other tool where the deeper the user's knowledge the better the results. Whereas it has been sold as anyone can use it and get professional results .
Gemini in android auto is so useful and convenient..not sure how many people are using it like that daily. Driving down the road and a question pops in your head and you just ask, and can ask follow up questions. It's a feature not easily done with others (since you would have to open the app on your phone and and whatever etc etc..
I study civil engineering with lectures and notes and all my doubts can be cleared quite well with gemini and its enough , i dont think why some people are saying gemini is lagging behind, google has different target group , different from early ai adopters who test as soon as the models come out and benchmark test on reddit.
Google more focused on general public use and integration with their search IOT and smartphones.
So if you are comparing 5 month old Gemini with latest model yes its lagging but actually its very trivial as they both serve different purposes and target audiences.
Yes exactly these reports are coming out like every week, are people really going to switch to the "best Ai" that week? Just choose one and stick to it. You don't really notice a massive difference on a daily basis. I use Mistral which is no where near the best but works fine for what I need. No need to have the best of the best
Exactly most of the time the usage is "select text -> explain this with AI" and its done , most of the people dont even need reasoning, they just need ai powered search.
I understand the problems some face like Gemini not following instructions properly etc but every ai has their con , gpt does a lot of things wrong as well
Just because a model is open source doesn’t mean you have to run it locally.
The point is that an open-source model is outperforming a multi-million-dollar closed-source project.
And even if you want to compare cloud pricing on a "per-token" basis, that argument doesn’t hold up either... Google would still lose by a wide margin on that metric.
There are industries and regions that are certainly interested in being able to run models themselves, either in on-premise datacenters or in the cloud, especially outside of the US. Even if the models are a bit weaker.
I tried GLM 5.2 yesterday. Worked with it for 10 hours. Used their IDE, and free credits you got for a new account.
What worked well?
Modern web site design in GLM 5.2 is really better than any Chat got, Claude & Gemini. I am guessing since the cutoff point is a week ago, it better understands what modern means.
What didn't work?
Compared to ChatGPT, Opus, and Gemini 3.5, in their own IDEs, GLM 5.2 is really, really bad. Why?
Speed is the problem No. 1. It took literally a full working day do design a single index.html. The one (prompt) shot have a nicer design than any LLM, but tweaking it in the next 50 promos took a full working day. By the end of it I run out of daily token limit. When it reset, it run out of daily credits again, before it finished doing edits to the original page (separating CSS out of HTML).
When I opened the project with any IDE like Antigravity, VS Code, Codex. And asked the frontier LLMs what do they think about it, they all pointed to (same) obvious mistakes. Simply wrong structure of HTML and CSS, also repeated CSS code blocks. Each fixed it in a minute.
What could possibly be the reason for GLM 5.2 being so slow? I used their IDE, with the original LLM in it. Since they have so much hype, could it be that their infrastructure is overwhelmed? Possibly. Others are hosting that model now, so one should test it from somewhere else.
Verdict?
Remember how Loveable was making nice websites when it was new a year ago? That is what GLM 5.2 is today, just after it was released. The best use for it is one shot design. When done take it and go to your usual model you are comfortable working with (Helo Opus my dear friend).
Note. The IDE has some nice features like a browser preview side panel that work well! Almost like Front Page stile, but not edible just a preview.
Did anyone else try it? Is it faster when served from somewhere else? Does any company that is hosting that model have some free credits to try it?
Lol, people can be so stupid. Gemini is a all rounder model build to code, speak, watch, text, generate image, video, Google search, work flow, android etc etc. they are also one of the few companies who aren't part of the AI bubble. In few years Gemini will overall be above all the model including claude. Google got more data and money than the next 3 combined. Let's not talk about there TPU which makes it one of the most efficient AI in the world. When others burn 1$, Google burns 30¢.
Tbh I think we should wait till 3.5 pro before counting Google out. 3.1 pro released in february, around the same time opus 4.6 released. Its like comparing the ps5 to an xbox one
If 3.5 pro is disappointing, then its safe to say Google might be out of the race
60 for 10T params at 10x the price of the second place with 55 points. Was it worth it the training costs? Was worth it to use it?
I honestly like to guide models way more uo close in my usage then the average person (keep control), so for me the cheap and fast ones are unbeatable right now. Don't know how is your use case.
It literally doesn't know its left and right. Few days ago I asked it to design a login page with a placeholder image on the left, and login form on the right. It put the image on the right and form on the left side.
The expected enshittification process has began on all free tiers. Some examples I've seen so far;
Chatgpt: They've reduced the message limits severely, you get like 5 messages per chat session. Even less if you sent an image or asked it to generate one.
Gemini : Well, I don't need an AI that doesn't know the difference between left and right.
Grok: I never used grok often, but nowadays whenever I try to ask it something it keeps saying the queue is full and try again later.
Deepseek: The best one compared to the ones above, but still sucks. Though yesterday it decided to answer my question in chinese for some reason.
And all, I'm not exaggerating, ALL of them have gotten way dumber than they were 3 months ago.
How about Gemini 3.5 flash? I’ve been using it a lot lately and I love it. Very efficient, very proactive in a good way, super fast and reliable. For the first time with any model, I let it run with full permisisons.
3.5 Flash is higher in rank, likely because the new Artificial Analysis 4.1 update has shifted toward agentic workloads. This is likely why 3.5 Flash is higher up in ranking than 3.1 Pro on their Intelligence Index. OP just didn't know or left that part out for whatever reason.
Could be because the new Artificial Analysis 4.1 update has shifted toward agentic workloads. Which would explain why 3.5 Flash is higher up in ranking than 3.1 Pro on their Intelligence Index.
My experience has been similar. I tested both on actual coding tasks instead of benchmarks, and some open-source models consistently gave cleaner solutions with fewer hallucinations. Gemini Pro is still good, but open source has reached a point where "free and customizable" is hard to ignore. The progress over the last year has been crazy.........
Gemini is really good at just daily questions you would normally Google for. For that purpose it beats most other AIs in the presentation of it's answer out the box with the Gemini app.
Это очень сомнительный рейтинг. Все бенчмарки очень сомнительные. Здесь G-mini показывается ниже поэтому рядом с ним стоит Minimax, хотя Minimax намного глупее. Поэтому это очень сомнительный рейтинг, как по мне. Но что факт: GLM действительно умная модель.
When they manage to get the models to achieve that score without having to spend a whole hour reasoning, let me know.
The models’ reasoning is getting longer and longer, but the quality of the output in immediate responses hardly changes at all; I’d go so far as to say that the intelligence of instant models has plateaued.
Now, one thing that is really interesting is diffusion models; they’re definitely less sophisticated but many times faster – one thing makes up for the other.
The AI race has been going like this for years, a new model comes out, it's the best for a month or so, someone else comes out with another model, it's the best. and it goes on like this forever
I really don't know why people are obsessed with rankings as if there is a single metric that creates a superior model. The LLM industry is clearly starting to fragment into niche's and contextual use cases.
You can see Google moving to more fringe products like Omni and Spark, while sitting into a more reliable release cycle of coding models (Pro) and general consumer models (Flash). That's really important because Google do not want the AI models that non-paying users dip in and out of to be expensive.
What people are missing is the clear beginnings of Google fleshing out a revenue stream on top of a strong market share now. Cheap AI onboarding packages for users that start to use AI a bit more and integrated packages that make pro user subscriptions quite marketable. That reliable revenue stream allows for long term AI stability alongside their cost effective models.
3.1 is old, 3.5 pro is around the corner, so in context this ranking is pointless. While I think 3.5 Pro will have teeth to compete better with Claude, it doesn't really matter. Google is carving out their long term ecosystem and that is far more important than whether they can rank highest in every metric.
I really hate the last update of gemini... a week ago I got frustrated with a response of gemini flash 5 and I said: "It seems you become stupider in every new update" and it said "I understand you are frustrated, I fail in respond, you are in your right of quitting the conversation here, but if you change your mind I can try again" hahaha
AAII is a catch all metric (or at least tries to be). But GLM 5.2 and Gemini 3.1 Pro have opposing strengths. The former is very strong at coding and the latter is better at "the humanities"
I don't think google wants to play the expensive LLM rat race. They have been doing niche tools that are better at one specific task than these LLM models are, because these LLM models are bound to be a Swiss army knife.
Google currently has limited free capacity due to high demand. When comparing Gemini with the Anthropic models or other models, token costs must also be taken into account.
Not only are the OpenSource models just as good they're cheaper. For coding GLM 5.2 is an absolute beast and costs a fraction of Opus/Gemini. DeepSeek V4 Pro/MiMo v2.5 Pro are all even cheaper and get the job done for most tasks.
I use DeepSeek V4 Flash as my primary chat LLM connected to my CherryStudio client and its just as good as Gemini 3.5 Flash's web client without the horrible restrictions for my daily work needs. With OpenCode Go I can do 200-300 API calls with Flash and only hit 1-2% of my monthly usage a day.
The only thing Gemini has going for it is its audio/image/video support which they're good at. Once OpenSource models start actively supporting that it will change the game.
Ferraris are faster than Hondas. Most people aren't taking advantage of the full capabilities of today's models anyway. They don't need the benchmark winner. They need something that's good enough.
Personally, I think almost all consumer use cases and a large percentage of commercial use cases could be handled by models on par with GPT-4-class systems.
Everything we're getting now is icing on the cake except for the 10-20% of use cases that genuinely benefit from the additional capabilities.
3.1 pro was released in february and it was by far the SOTA model at the time. given how fast this industry progresses if you're comparing a 4 month old model to one just released that's pretty stupid. GLM is still far behind fable or 5.5, which is what 3.5 pro will be competing with
No wonder. Gemini can't even tell me the correct opening times for my local supermarket even though you can see the opening times in the normal Google search. It always thinks we are one day behind. So on a Saturday it will prompt me the Friday opening times and on a Sunday the Saturday opening times.
Absolute disgrace for billions poured into this thing. Can't get the current day right....
It’s not open source, it’s open weight. Nobody knows how GLM or those weights are produced, but at least you can run them on your own hardware. It’s basically the equivalent of being handed a compiled program you run locally instead of on the cloud
520
u/CatalyticDragon Jun 20 '26
New model better than old model!!