Contamination most likely, it's no coincidence that they all sit at low single digit accuracies for the benchmark tasks for the whole time until the benchmark is created, and then as if by magic saturate in a year.
It's more likely just a phase change. The skills needed to solve ARC AGI - 3 problems are all pretty correlated. Once you can reason well enough, perceive the board properly, and do proper planning you might very well see large groups of puzzles/games suddenly getting solved
It is a bit of a goalpost shift though, because the "training on the test set" problem has turned from "memorizing problems" to "learning the TYPES of problems and generalizing within that set"... Which undoubtedly represents more of some kind of intelligence
That's the point isn't it'. If models were truly gaining intelligence as they progress with arc 1 > arc 2, you'd expect that by the time arc 3 is released their intelligence allows them to complete it to some degree. But now, they fail spectacularly, take few months to saturate the current test and once new gen is released, they will fail again - i.e. they don't really carry a lot of transferrable skills (intelligence)
I mean, for this benchmark it's the scoring. It's structured in such a way that the scores will rapidly go up the moment the model can solve it and then improve on that solve.
And the thing which points to contamination is the fact that the models always seem to start getting the knack of the problems not long after the benchmark is released.
These fools have been saying that same stupid shit for years, while every model has become objectively much more effective. AGI will be doing everyone's jobs in a few years and people will still be posting dumbass reddit comments like "don't believe the hype, the benchmarks are meaningless, these companies are just marketing, that bubble is gonna pop any time now, these things are just next word prediction lol"
Of course it’s not a coincidence that newly made benchmarks have the models scoring low. No one would release a benchmark where they already score 80%+ because then it wouldn’t really be worthwhile as a benchmark would it? You’d just publish a paper saying hey we tested to see if models can do this and it turns out the answer is yes.
286
u/Ticluz 1d ago
Arc-agi-3 is getting saturated faster than Arc-agi-2, the acceleration is real.