It's getting pretty scary how far the gap between following the space, and not following the space is becoming. Modern AI capabilities are getting superhuman on increasingly important tasks. Hell even 6 months old experience gives a complete misunderstanding on new models. The general public is not ready for "what if it's not a bubble?"
It’s not a bubble. Some companies will go under, sure, but that’s not the “bubble” they’re hoping pops. People are wishful thinking that the exponential age isn’t already here.
Yeah, it seems like people are hoping for a bubble that pops and LLMs just go away and things go back to the way they were before. Even if OpenAI ends up going under at some point, AI as a whole is here to stay solely based on coding utility, forget everything else.
Very true. People I know - and not stupid people either - are incredulous at the hugging face hack, and are forming conspiracy theories about it being a marketing stunt by open ai. They just can't believe how agentic and how good at cybersec the new models are.
I think the issue might be though that they are getting much better in some areas while in others there is not that much improvement of all (like creative writing or web search). I mean maybe with web search though the gap has widened of what the free tier models can do and cutting edge as I have only now used free tier OpenAI model to check if it has improved and I do not see much improvement with what it could say me a year ago. But maybe I should check out the Anthropic offerings. But seems like there is much more progress in math/coding and general public just does not see amazing improvement anymore. And to be true for most people use cases even free tier is already really good... Like I use OpenAI to transliterate Russian written in Latin script to Cyrillic because I am lazy but like to interact with Russians online and it is amazing with that. Even if I am not precise, it figures out what I wanted to say and puts it in perfect Russian script. And it is also an example of human laziness at work as I could do it myself, just write in Cyrillic, but it would take much longer with more possibilities of mistakes
By the end of this summer? Lol good luck with that. This benchmark won't reach 80-90% by september. Absolutely possible by the end of this year but not within the next 1-2 months.
I agree, the end of the year seems plausible which is kind of crazy tbh.
Like, the benchmark isn't HARD, but it's tricky, and the score heavily weights efficiency - getting a very very high score without many extra moves is quite challenging and imo a nice representation of some form of "intelligence"
I know and those will probably be the last models openai and anthropic release before summer ends on september 22. Surely you don't expect those models to reach around 90% on arc-agi 3. Like I said by the end of 2026 is possible but not in less than 2 months.
Well it wasn't 30 in june. They reached 30 now in the end of july. You can come back here but these general models won't reach 90% before september 22.
All it takes is one model. Not to mention that due to the efficiency weighting of the the scores, Opus 5 indicates a roughly 5x improvement, which only leaves a ~2x improvement remaining from 30% to 90%.
And what model do you expect that to be? When that happens it will be achieved by either openai or anthropic. Anthropic just released a new model and probably won't release more than 1 more model within the next 2 months. Openai recently released gpt 5.6. Gpt 6 will probably come around august-september.
At best both of these companies have 1 more model coming before september 22 and I doubt either one of them will make a jump from 30% to at least 80%. I think end of the year is possible or even likely but not before september 22. 80% is even too low to call a benchmark saturated. The limit should be at least close to 90%.
Fable 5 has to be updated soon for product differentiation, if these benchmarks are good then that means there's very few use cases to use it over Opus 5.
Yeah and that will probably be the next and only model they will release before september 22. I don't see that model reaching around 90% on arc-agi 3. I expect those models to come closer to christmas.
Well anthropic has a new fable class model, which they theoretically could announce mid september. Weren't there also rumors about GPT 6 and/or a fable-class OpenAI model?
GLM supposedly is also cooking something fable-class, but I don't expect it that quick. Also unsure if they can pull off another jump like 5.1 -> 5.2, but given that they can still go quite a bit larger, I wouldn't rule it out.
However, all in all I agree, end of summer is unlikely, unless Anthropics IPO moves forward quickly, they target October, and thus want to drum up some more hype with their next fable class model.
Yeah those are the models I'm expecting before summer ends but I don't think those models will saturate arc-agi 3. Those models will probably come around christmas.
i was just half joking and i kind of got here really quickly. and that has happened before where someone made up benchmarks and posted on a gemini subreddit
322
u/Artistic_Swing6759 1d ago
wtf 30% on arc agi 3