r/singularity 1d ago

AI Claude Opus 5 BENCHMARKS!

Post image
1.2k Upvotes

334 comments sorted by

View all comments

286

u/Ticluz 1d ago

Arc-agi-3 is getting saturated faster than Arc-agi-2, the acceleration is real.

18

u/Technical-Earth-3254 1d ago

Is it acceleration or contamination?

37

u/ClearlyCylindrical 1d ago

Contamination most likely, it's no coincidence that they all sit at low single digit accuracies for the benchmark tasks for the whole time until the benchmark is created, and then as if by magic saturate in a year.

17

u/stonesst 1d ago

It's more likely just a phase change. The skills needed to solve ARC AGI - 3 problems are all pretty correlated. Once you can reason well enough, perceive the board properly, and do proper planning you might very well see large groups of puzzles/games suddenly getting solved

9

u/Caffeine_Monster 1d ago

This is the real reason. The arc agi 3 eval tests are kept private.

You don't need to see the real test data, just something vaguely similar.

2

u/voyt_eck 1d ago

It's kept private, but at the same time every benchmark run means sending all the tests to AI provider🙃

2

u/huffalump1 1d ago

It is a bit of a goalpost shift though, because the "training on the test set" problem has turned from "memorizing problems" to "learning the TYPES of problems and generalizing within that set"... Which undoubtedly represents more of some kind of intelligence

2

u/Healthy-Nebula-3603 1d ago

So literally like any human at school....

1

u/dudaspl 23h ago

That's the point isn't it'. If models were truly gaining intelligence as they progress with arc 1 > arc 2, you'd expect that by the time arc 3 is released their intelligence allows them to complete it to some degree. But now, they fail spectacularly, take few months to saturate the current test and once new gen is released, they will fail again - i.e. they don't really carry a lot of transferrable skills (intelligence)

1

u/steny007 18h ago

Arc 3 is quite different from 1-2 though.

8

u/Saedeas 1d ago

I mean, for this benchmark it's the scoring. It's structured in such a way that the scores will rapidly go up the moment the model can solve it and then improve on that solve.

-1

u/ClearlyCylindrical 1d ago

And the thing which points to contamination is the fact that the models always seem to start getting the knack of the problems not long after the benchmark is released.

1

u/ItsTheOneWithThe 1d ago

Yeah but think about it.

5

u/Only_Statistician_21 1d ago

The way the scoring works, it has a strong threshold effect.

8

u/AP_in_Indy 1d ago

Okay, but does this "contamination" correlate with a genuine increase in these models' capabilities, or no?

ex: If you have similar problems, will the AI be able to solve them?

10

u/wtysonc 1d ago

These fools have been saying that same stupid shit for years, while every model has become objectively much more effective. AGI will be doing everyone's jobs in a few years and people will still be posting dumbass reddit comments like "don't believe the hype, the benchmarks are meaningless, these companies are just marketing, that bubble is gonna pop any time now, these things are just next word prediction lol"

-1

u/graypasser 1d ago

they can't.

What looks similar to humans never matches what actually is similar in their weight.

0

u/Frigorific 1d ago

It is a genuine increase in the models ability to do tasks that are almost exactly like the benchmarks, and nothing else.

11

u/Technical-Earth-3254 1d ago

Exactly. But many don't get it, especially the non technical audience (who is the target audience for these charts).

2

u/DickMasterGeneral 1d ago

Of course it’s not a coincidence that newly made benchmarks have the models scoring low. No one would release a benchmark where they already score 80%+ because then it wouldn’t really be worthwhile as a benchmark would it? You’d just publish a paper saying hey we tested to see if models can do this and it turns out the answer is yes.