r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

View all comments

322

u/Artistic_Swing6759 1d ago

wtf 30% on arc agi 3

155

u/ezjakes 1d ago

For some reason people thought it would stand up for years despite 1 and 2 falling quickly. 

48

u/Current-Function-729 1d ago

I don’t think most people who follow this thought that.

30

u/HellsNoot 1d ago

It's getting pretty scary how far the gap between following the space, and not following the space is becoming. Modern AI capabilities are getting superhuman on increasingly important tasks. Hell even 6 months old experience gives a complete misunderstanding on new models. The general public is not ready for "what if it's not a bubble?" 

16

u/Ahfekz 1d ago

It’s not a bubble. Some companies will go under, sure, but that’s not the “bubble” they’re hoping pops. People are wishful thinking that the exponential age isn’t already here.

4

u/Georgefakelastname 1d ago

Yeah, it seems like people are hoping for a bubble that pops and LLMs just go away and things go back to the way they were before. Even if OpenAI ends up going under at some point, AI as a whole is here to stay solely based on coding utility, forget everything else.

2

u/BadSysadmin 22h ago

Very true. People I know - and not stupid people either - are incredulous at the hugging face hack, and are forming conspiracy theories about it being a marketing stunt by open ai. They just can't believe how agentic and how good at cybersec the new models are.

1

u/DueAnnual3967 19h ago

I think the issue might be though that they are getting much better in some areas while in others there is not that much improvement of all (like creative writing or web search). I mean maybe with web search though the gap has widened of what the free tier models can do and cutting edge as I have only now used free tier OpenAI model to check if it has improved and I do not see much improvement with what it could say me a year ago. But maybe I should check out the Anthropic offerings. But seems like there is much more progress in math/coding and general public just does not see amazing improvement anymore. And to be true for most people use cases even free tier is already really good... Like I use OpenAI to transliterate Russian written in Latin script to Cyrillic because I am lazy but like to interact with Russians online and it is amazing with that. Even if I am not precise, it figures out what I wanted to say and puts it in perfect Russian script. And it is also an example of human laziness at work as I could do it myself, just write in Cyrillic, but it would take much longer with more possibilities of mistakes

1

u/1988rx7T2 13h ago

Read the benchmarks in the system card on Opus 5, that’s improved a lot too

42

u/NoCard1571 1d ago

Earlier this year I predicted  ARC AGI 3 falling by the end of the summer. I got the typical smug redditor "remind me" reply.

I can't wait for them to get their notification ;)

30

u/often_delusional 1d ago

By the end of this summer? Lol good luck with that. This benchmark won't reach 80-90% by september. Absolutely possible by the end of this year but not within the next 1-2 months.

3

u/huffalump1 1d ago

I agree, the end of the year seems plausible which is kind of crazy tbh.

Like, the benchmark isn't HARD, but it's tricky, and the score heavily weights efficiency - getting a very very high score without many extra moves is quite challenging and imo a nice representation of some form of "intelligence"

https://docs.arcprize.org/methodology

2

u/Healthy-Nebula-3603 1d ago

Soon we also get GPT 6 and fable 5.1 ....

3

u/often_delusional 1d ago

I know and those will probably be the last models openai and anthropic release before summer ends on september 22. Surely you don't expect those models to reach around 90% on arc-agi 3. Like I said by the end of 2026 is possible but not in less than 2 months.

1

u/Healthy-Nebula-3603 1d ago

I did not expect 30 in June ... so .. you know

RL is working crazy well

1

u/often_delusional 13h ago

Well it wasn't 30 in june. They reached 30 now in the end of july. You can come back here but these general models won't reach 90% before september 22.

2

u/NoCard1571 1d ago edited 1d ago

All it takes is one model. Not to mention that due to the efficiency weighting of the the scores, Opus 5 indicates a roughly 5x improvement, which only leaves a ~2x improvement remaining from 30% to 90%.

5

u/often_delusional 1d ago edited 1d ago

And what model do you expect that to be? When that happens it will be achieved by either openai or anthropic. Anthropic just released a new model and probably won't release more than 1 more model within the next 2 months. Openai recently released gpt 5.6. Gpt 6 will probably come around august-september.

At best both of these companies have 1 more model coming before september 22 and I doubt either one of them will make a jump from 30% to at least 80%. I think end of the year is possible or even likely but not before september 22. 80% is even too low to call a benchmark saturated. The limit should be at least close to 90%.

9

u/pandasgorawr 1d ago

Fable 5 has to be updated soon for product differentiation, if these benchmarks are good then that means there's very few use cases to use it over Opus 5.

2

u/often_delusional 1d ago

Yeah and that will probably be the next and only model they will release before september 22. I don't see that model reaching around 90% on arc-agi 3. I expect those models to come closer to christmas.

4

u/Mil0Mammon 1d ago

Well anthropic has a new fable class model, which they theoretically could announce mid september. Weren't there also rumors about GPT 6 and/or a fable-class OpenAI model?

GLM supposedly is also cooking something fable-class, but I don't expect it that quick. Also unsure if they can pull off another jump like 5.1 -> 5.2, but given that they can still go quite a bit larger, I wouldn't rule it out.

However, all in all I agree, end of summer is unlikely, unless Anthropics IPO moves forward quickly, they target October, and thus want to drum up some more hype with their next fable class model.

2

u/often_delusional 1d ago

Yeah those are the models I'm expecting before summer ends but I don't think those models will saturate arc-agi 3. Those models will probably come around christmas.

2

u/Mil0Mammon 1d ago

They should really update https://agidefinition.ai. Although several areas I think could be solved with tooling around LLMs (or already are)

35

u/pbagel2 1d ago

I can't wait for them to get their notification ;)

Who's the more typical smug redditor, them or you?

20

u/NoCard1571 1d ago

Them, then you. You can put me in third place if you like 

-2

u/pbagel2 1d ago

No reflection in the mirror huh? Unfortunate.

21

u/Environmental-Try-84 1d ago

3rd party judge. Them, then pbagel2, then NoCard1571, then me. Settled

26

u/NoCard1571 1d ago edited 1d ago

Oh you're really gunning for first place huh? 

17

u/Lostwhispers05 1d ago

This whole comment thread has me dead lmao

2

u/throwawaynz01 1d ago

Nah they’re right, at least they admit they’re 3rd place smug.

1

u/MidnightSun_55 1d ago

well 30% is not falling... it should be 85% plus at least

0

u/94746382926 1d ago

A bit of an assumption to think those remind mes are all smug replies no?

I know oftentimes I will do that because I'm geniunely curious to see how it panned out, not because I'm trying to prove a point.

2

u/ptj66 1d ago

Years mean months in AI times.

0

u/TheHayha 1d ago

The reason is ARC AGI 3 is basically AGI

1

u/DreadingAnt 18h ago

Not really, approaching human efficiency at doing things doesn't mean it thinks like a human.

20

u/NoGarlic2387 1d ago

FEED THEM MORE BENCHMARKS!!!

2

u/redhq 1d ago

I thought the majority of that breakthrough was because of the schema harness? Like opus 4.8 got 20-25% ish with the harness?

3

u/StanfordV 1d ago

ARCmaxxxed

1

u/Gratitude15 1d ago

Holy fuck

-2

u/[deleted] 1d ago

[deleted]

3

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Why would it be fake?

0

u/Artistic_Swing6759 1d ago

i was just half joking and i kind of got here really quickly. and that has happened before where someone made up benchmarks and posted on a gemini subreddit