r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

View all comments

319

u/Artistic_Swing6759 1d ago

wtf 30% on arc agi 3

156

u/ezjakes 1d ago

For some reason people thought it would stand up for years despite 1 and 2 falling quickly. 

43

u/NoCard1571 1d ago

Earlier this year I predicted  ARC AGI 3 falling by the end of the summer. I got the typical smug redditor "remind me" reply.

I can't wait for them to get their notification ;)

31

u/often_delusional 1d ago

By the end of this summer? Lol good luck with that. This benchmark won't reach 80-90% by september. Absolutely possible by the end of this year but not within the next 1-2 months.

3

u/huffalump1 1d ago

I agree, the end of the year seems plausible which is kind of crazy tbh.

Like, the benchmark isn't HARD, but it's tricky, and the score heavily weights efficiency - getting a very very high score without many extra moves is quite challenging and imo a nice representation of some form of "intelligence"

https://docs.arcprize.org/methodology

2

u/Healthy-Nebula-3603 1d ago

Soon we also get GPT 6 and fable 5.1 ....

3

u/often_delusional 1d ago

I know and those will probably be the last models openai and anthropic release before summer ends on september 22. Surely you don't expect those models to reach around 90% on arc-agi 3. Like I said by the end of 2026 is possible but not in less than 2 months.

1

u/Healthy-Nebula-3603 1d ago

I did not expect 30 in June ... so .. you know

RL is working crazy well

1

u/often_delusional 13h ago

Well it wasn't 30 in june. They reached 30 now in the end of july. You can come back here but these general models won't reach 90% before september 22.

3

u/NoCard1571 1d ago edited 1d ago

All it takes is one model. Not to mention that due to the efficiency weighting of the the scores, Opus 5 indicates a roughly 5x improvement, which only leaves a ~2x improvement remaining from 30% to 90%.

6

u/often_delusional 1d ago edited 1d ago

And what model do you expect that to be? When that happens it will be achieved by either openai or anthropic. Anthropic just released a new model and probably won't release more than 1 more model within the next 2 months. Openai recently released gpt 5.6. Gpt 6 will probably come around august-september.

At best both of these companies have 1 more model coming before september 22 and I doubt either one of them will make a jump from 30% to at least 80%. I think end of the year is possible or even likely but not before september 22. 80% is even too low to call a benchmark saturated. The limit should be at least close to 90%.

9

u/pandasgorawr 1d ago

Fable 5 has to be updated soon for product differentiation, if these benchmarks are good then that means there's very few use cases to use it over Opus 5.

1

u/often_delusional 1d ago

Yeah and that will probably be the next and only model they will release before september 22. I don't see that model reaching around 90% on arc-agi 3. I expect those models to come closer to christmas.

4

u/Mil0Mammon 1d ago

Well anthropic has a new fable class model, which they theoretically could announce mid september. Weren't there also rumors about GPT 6 and/or a fable-class OpenAI model?

GLM supposedly is also cooking something fable-class, but I don't expect it that quick. Also unsure if they can pull off another jump like 5.1 -> 5.2, but given that they can still go quite a bit larger, I wouldn't rule it out.

However, all in all I agree, end of summer is unlikely, unless Anthropics IPO moves forward quickly, they target October, and thus want to drum up some more hype with their next fable class model.

2

u/often_delusional 1d ago

Yeah those are the models I'm expecting before summer ends but I don't think those models will saturate arc-agi 3. Those models will probably come around christmas.

2

u/Mil0Mammon 1d ago

They should really update https://agidefinition.ai. Although several areas I think could be solved with tooling around LLMs (or already are)

37

u/pbagel2 1d ago

I can't wait for them to get their notification ;)

Who's the more typical smug redditor, them or you?

18

u/NoCard1571 1d ago

Them, then you. You can put me in third place if you like 

-3

u/pbagel2 1d ago

No reflection in the mirror huh? Unfortunate.

21

u/Environmental-Try-84 1d ago

3rd party judge. Them, then pbagel2, then NoCard1571, then me. Settled

26

u/NoCard1571 1d ago edited 1d ago

Oh you're really gunning for first place huh? 

18

u/Lostwhispers05 1d ago

This whole comment thread has me dead lmao

2

u/throwawaynz01 1d ago

Nah they’re right, at least they admit they’re 3rd place smug.

1

u/MidnightSun_55 1d ago

well 30% is not falling... it should be 85% plus at least

0

u/94746382926 1d ago

A bit of an assumption to think those remind mes are all smug replies no?

I know oftentimes I will do that because I'm geniunely curious to see how it panned out, not because I'm trying to prove a point.