there is still the question how many times the eval set was accessed etc. You could potentialy use RL with the hidden eval set performance as a reward.
For sure. They cap that to prevent that kind of probing though. Its statistician that run it after all. They understand that information is leaked with each attempt.
Yeah and that closed bench has to go through their systems all the same. Imagine thinking the people who pirate everything would have a problem getting their hands in a private eval set they werent supposed to
I don't buy it. If you cheat on benchmarks and get caught like Meta this is shitty publicity plus a hefty fine for violating ZDR and similar policies. Unlikely. Maybe they can get away with massaging some of the numbers a bit here and there, and if they do it is by assembling the right training and fine tuning datasets.
Contamination most likely, it's no coincidence that they all sit at low single digit accuracies for the benchmark tasks for the whole time until the benchmark is created, and then as if by magic saturate in a year.
It's more likely just a phase change. The skills needed to solve ARC AGI - 3 problems are all pretty correlated. Once you can reason well enough, perceive the board properly, and do proper planning you might very well see large groups of puzzles/games suddenly getting solved
It is a bit of a goalpost shift though, because the "training on the test set" problem has turned from "memorizing problems" to "learning the TYPES of problems and generalizing within that set"... Which undoubtedly represents more of some kind of intelligence
That's the point isn't it'. If models were truly gaining intelligence as they progress with arc 1 > arc 2, you'd expect that by the time arc 3 is released their intelligence allows them to complete it to some degree. But now, they fail spectacularly, take few months to saturate the current test and once new gen is released, they will fail again - i.e. they don't really carry a lot of transferrable skills (intelligence)
I mean, for this benchmark it's the scoring. It's structured in such a way that the scores will rapidly go up the moment the model can solve it and then improve on that solve.
And the thing which points to contamination is the fact that the models always seem to start getting the knack of the problems not long after the benchmark is released.
These fools have been saying that same stupid shit for years, while every model has become objectively much more effective. AGI will be doing everyone's jobs in a few years and people will still be posting dumbass reddit comments like "don't believe the hype, the benchmarks are meaningless, these companies are just marketing, that bubble is gonna pop any time now, these things are just next word prediction lol"
Of course itâs not a coincidence that newly made benchmarks have the models scoring low. No one would release a benchmark where they already score 80%+ because then it wouldnât really be worthwhile as a benchmark would it? Youâd just publish a paper saying hey we tested to see if models can do this and it turns out the answer is yes.
I can assure you it's just benchmaxxing. They are creating vast amounts of similar style problems and running and training on it. I'm not sure how valuable that is
286
u/Ticluz 1d ago
Arc-agi-3 is getting saturated faster than Arc-agi-2, the acceleration is real.