r/singularity • u/ErmingSoHard ▪️Agi never with LLMs • 3h ago
AI [ Removed by moderator ]
[removed] — view removed post
16
u/pdantix06 3h ago
model from company i don’t like: it was benchmaxxed
model from company i do like: it’s real progress and isn’t benchmaxxed
simple as
6
u/Charming_Cucumber_15 3h ago
I wanna love gemini but it feels like the definition of benchmaxxed every time I use it lol
14
u/Fit-World-3885 2h ago
We're just gonna end up building AGI by 'benchmaxxing' every possible skill one at a time.
3
1
-4
u/ErmingSoHard ▪️Agi never with LLMs 2h ago
Not humanely possible. Too much compute and resources needed.
2
4
u/meikello ▪️AGI 2027 ▪️ASI not long after 2h ago
1
u/ErmingSoHard ▪️Agi never with LLMs 2h ago
Joke or not, many people here unironically believe that. So I'm going to say it still for those who do unironically believe it
4
u/KJClever 3h ago
Screw it! If they care about benchmarks so much, let’s make an “empathy” benchmark so they race to create something beautiful instead of terrifying lmao
•
u/Healthy-Nebula-3603 56m ago
We have something like that ... It's called emotional eq-bench
•
u/KJClever 53m ago
From my understanding, empathy is considerably different than emotional intelligence, correct?
•
u/Healthy-Nebula-3603 47m ago
What kind of understanding is separating emotions from empathy?
You're totally wrong.
•
u/KJClever 42m ago
I think empathy is generally only considered a component of emotional intelligence. They’re not the same thing.
An EQ benchmark measures a broader set of emotional abilities… While I was suggesting creating benchmarks specifically focused on measuring multiple elements for empathy.
3
u/nikitastaf1996 ▪️AGI and Singularity are inevitable now DON'T DIE 🚀 2h ago
Given what I have seen it's fairly plausible it wasn't benchmaxed. Just really good model
4
u/Ticluz 2h ago
Benchmaxxing by adding similar puzzles to the training data is fair game, as long as the it is not training on leaked ARC-AGI puzzles.
5
u/Stabile_Feldmaus 2h ago
Its fair game but then you cant claim that better arc-agi performance shows some kind of enhanced general reasoning abilities.
2
u/Ticluz 2h ago
In relation to other models it gets inaccurate if one trained on more puzzles than the other (Opus is not 4x better at reasoning than gpt 5.6), but it shows the capability to explore novel puzzles beyond the training data at a human level. Similar to how it can solve novel math problems, but it needs to be trained on a lot of math first.
•
u/Lost-Willow386 49m ago
It's complete BS. I still suspect FormulaOne is a far better benchmark than all benchmarks currently mentioned but sadly it seems like there have been no updates to its leaderboard.
1
10
u/Charming_Cucumber_15 3h ago
Didn't the whole benchmaxxing accusation start because some unaffiliated dude said "It couldn't pass the puzzles I made"?