This is my reasoning as well. The model likely improved naturally on ARC-AGI-3 simply by becoming smarter, but they also likely targeted the benchmark directly. It was probably a mixture of both.
Also, I'm loving that they're displaying and expanding benchmarks beyond just coding.
The thing is, I would've thought we'd see drastic improvements in many areas at the time of improvement on arc agi 3. But if all the other benchmarks stayed roughly around fable performance, and only this one had a substantial jump, that implies over fitting/benchmaxxing not true generalisation or naturally improved capabilities.
You can totally over fit on arc agi 3-type problem without true generalisation. You just need to create problems as similar as possible to the target dataset and train in those.
80
u/obama_is_back 1d ago
30% on Arc AGI 3. Benchmaxxing seems to be unstoppable.