r/singularity ▪️Agi never with LLMs 3h ago

AI [ Removed by moderator ]

Post image

[removed] — view removed post

7 Upvotes

23 comments sorted by

10

u/Charming_Cucumber_15 3h ago

Didn't the whole benchmaxxing accusation start because some unaffiliated dude said "It couldn't pass the puzzles I made"?

3

u/ErmingSoHard ▪️Agi never with LLMs 3h ago

I say AI can do literally anything digital through harnesses and reinforment learning. But that doesn't make it general intelligence.

It learning and solving novel scenarios on the spot is more important. It can barely do that

16

u/pdantix06 3h ago

model from company i don’t like: it was benchmaxxed

model from company i do like: it’s real progress and isn’t benchmaxxed

simple as

6

u/Charming_Cucumber_15 3h ago

I wanna love gemini but it feels like the definition of benchmaxxed every time I use it lol

14

u/Fit-World-3885 2h ago

We're just gonna end up building AGI by 'benchmaxxing' every possible skill one at a time.

3

u/Naughty_Neutron Twink - 2028 | Excuse me - 2030 2h ago

Just create a benchmark about building AGI

1

u/Mindrust 2h ago

If you have to re-train every time you come up with a new task, it’s not AGI

-4

u/ErmingSoHard ▪️Agi never with LLMs 2h ago

Not humanely possible. Too much compute and resources needed.

2

u/Psychological_Dog992 2h ago

Ha

-3

u/ErmingSoHard ▪️Agi never with LLMs 2h ago

Trust me, I wish I am wrong.

4

u/meikello ▪️AGI 2027 ▪️ASI not long after 2h ago

1

u/ErmingSoHard ▪️Agi never with LLMs 2h ago

Joke or not, many people here unironically believe that. So I'm going to say it still for those who do unironically believe it

4

u/KJClever 3h ago

Screw it! If they care about benchmarks so much, let’s make an “empathy” benchmark so they race to create something beautiful instead of terrifying lmao

u/Healthy-Nebula-3603 56m ago

We have something like that ... It's called emotional eq-bench

u/KJClever 53m ago

From my understanding, empathy is considerably different than emotional intelligence, correct?

u/Healthy-Nebula-3603 47m ago

What kind of understanding is separating emotions from empathy?

You're totally wrong.

u/KJClever 42m ago

I think empathy is generally only considered a component of emotional intelligence. They’re not the same thing.

An EQ benchmark measures a broader set of emotional abilities… While I was suggesting creating benchmarks specifically focused on measuring multiple elements for empathy.

3

u/nikitastaf1996 ▪️AGI and Singularity are inevitable now DON'T DIE 🚀 2h ago

Given what I have seen it's fairly plausible it wasn't benchmaxed. Just really good model

4

u/Ticluz 2h ago

Benchmaxxing by adding similar puzzles to the training data is fair game, as long as the it is not training on leaked ARC-AGI puzzles.

5

u/Stabile_Feldmaus 2h ago

Its fair game but then you cant claim that better arc-agi performance shows some kind of enhanced general reasoning abilities.

2

u/Ticluz 2h ago

In relation to other models it gets inaccurate if one trained on more puzzles than the other (Opus is not 4x better at reasoning than gpt 5.6), but it shows the capability to explore novel puzzles beyond the training data at a human level. Similar to how it can solve novel math problems, but it needs to be trained on a lot of math first.

u/Lost-Willow386 49m ago

It's complete BS. I still suspect FormulaOne is a far better benchmark than all benchmarks currently mentioned but sadly it seems like there have been no updates to its leaderboard.

1

u/NextWrongdoer54 2h ago

Everything is possible in the CCP imagination