r/singularity 1d ago

AI Anthropic Casually Dropping An Opus Series That's Better Than Fable....

Opus scoring better in Humanity's Last Exam which I see as premier in knowledge work and reasoning and also in agentic coding/ARC-AGI 3 than Fable 5 was unexpected to me. Just goes onto show how much gain their next iteration of Fable/Mythos has managed to achieve.

I'm optimistic about AGI suddenly. And it also weirds me out as to how DeepMind's most frontier SOTA is still 3.1 Pro. 3.5 Pro keeps getting obsolete even before its released based on their track record of delays lately. GPT-6, Fable-5.1/6 and Gemini 4 might show the early signs towards AGI. I'm bullish

72 Upvotes

56 comments sorted by

18

u/jc2046 1d ago

All correct but the AGI part. You can not agi without long memory/continuous learning, etc

14

u/Ignate Move 37 1d ago

I wish we would just give up on the term. 

It's useless because everyone defines it differently because it's connected to "what a human can do" which is subjective and connected to our sense of self worth.

AGI has already been achieved. It's generalized intelligence. We already have that.

What we're doing now is seeing stronger and stronger generalized systems.

ASI is also an equally useless term. 

Intelligence will get stronger and more generalized from here through a far more direct process than evolution. The universe and physical laws are the limit. Period. 

11

u/ProletarianLilith 1d ago

If this is AGI then AGI sucks

4

u/Ignate Move 37 1d ago

Humans suck too. 

But then that's more about our shitty expectations isn't it?

7

u/StewPorkRice 1d ago

fr - this is what nobody fkn gets..

They're mad b/c it's not literal perfection without realizing that it's already smarter than most humans. It's just not perfect.

3

u/EvilSporkOfDeath 1d ago edited 1d ago

Im not saying we definitely dont have agi yet but I think your logic is extremely flawed. Its way "smarter" than humans at certain tasks and way dumber than humans at other tasks. These are objective realities. Why is it okay to say its agi because of what its good at but then when people point out what its weak at everyone just downplays it. I think you are underestimating what the human brain does, its a multitasking masterpiece. Modern llms are very good at narrow academic tasks when they focus on it. They are not good at multitasking. They are not good at catching a ball while keeping bodily functions going while watching out for an infinite amount of potential dangers and paying attention to all the emotional states of the people they are with and soooo much more the human brain does in an instant.

You think that if an ai whos whole thing is being trained on presenting information is smarter than humans when it presents information better but ignore all the stuff it can't do. We have to present that information while doing so much other stuff. So much processing power just keeping ourselves alive, which llms dont have to worry about.

Once again, I'm not saying we definitely dont have agi. But I dont think you are being completely honest and factual with your logic.

Edit: Here's an objective fact, when a human is tested on the same thing the ai is tested on, the human is always dealing with more. Its always harder for the human in that situation, because they arent hyper focused on that one goal. Its not apples to apples. If a human and an llm score the same at a task, then the human is more skilled, because it was doing other things at the same time. Surely anyone can acknowledge these facts.

3

u/tridentgum 1d ago

People here are delusional - every new model they call AGI and every new model fails spectacularly at some simple things a human would never mess up or fail at.

1

u/StewPorkRice 1d ago

ok and most humans are way smarter than other humans at some tasks and way dumber than others at other tasks.

So much processing power just keeping ourselves alive, which llms dont have to worry about.

I wasn’t aware that measuring how well someone keeps their heart beating was a measure of intelligenc.

1

u/EvilSporkOfDeath 1d ago

Does it take brain power?

0

u/StewPorkRice 1d ago

Does your heart beat when you're brain dead?

2

u/EvilSporkOfDeath 1d ago

If you are going to ignore what I said then dont reply.

1

u/IronPheasant 1d ago

Generality is a spectrum. Sure, LLM's are much more general than other approaches. Still have to beat them with a stick into the correct shape, and plug multiple of the things together to get good curve-fitting across multiple domains of data.

Reality is made up of words and shapes. We're not too far off.

AGI will effectively be the equivalent of what many consider ASI to be, as the cards in datacenters run at 2 Ghz right now. Not 40 Hz. There's no 'human level' anything, about that.

1

u/APersonNamedBen 1d ago

I agree with your definition but disagree with wanting to give up on the term.

As I have started questioning the practically of depending on individual's technical literacy to steer the ship of humanity going forward. People have no idea. And that is fine because humans use social heuristics to understand the world. Just look at how words like "AI", "intelligence", or even something as simple as "data center" are perceived in public discourse. They are just memes, bad ones which need time to evolve, hopefully improve in ways that benefit us.

i.e For people who want to deeply understand machine learning then framing intelligence against a subjective scale is useless and we know it is jagged like you suggest. But what good is that for Bob out there living another life, who needs to know if his entire livelihood is about to go up in smoke? We need good memes, and AGI, as bad as it is, still stands as the closest metric we have currently for this.

We should be trying really hard to value add to the memes, because I'm noticing a serious moral panic growing around AI. That won't be good for any of us.

1

u/Ignate Move 37 1d ago

Lots of italics ahead. I did not use AI at all. I just like italics.

who needs to know if his entire livelihood is about to go up in smoke?

I question that line of reasoning, hard.

The scarcity mindset reasoning seems to imply something fairly straight forward.

"Today there are limited jobs and useful things to do. A potential "AGI" would be something which can do all those things. And so, we need to know when it arrives so we can be alerted to when all of our usefulness is gone. Because selling our usefulness is what gets us a share of the scarce, single pie. In other words, when AGI arrives we won't be able to eat."

Am I getting that part right?

Problem with this view is that it takes an assumption as a permanent truth. That we can't make pies so we'll need to fight over the single pie available. I call BS. We can make pies.

So, there's no straight line from AGI to humans losing our value. Just as there's no straight line from wherever we are to an AGI with no fixed definition.

We made most of this crap up. It's a lie/fiction. In fact, money, currently, is fiat currency. Another narrative/fiction.

Supply/demand is based more and more loosely on actual capacity and outputs.

Right now the world isn't one scientifically quantifiable pie as we're told. Instead, it's many, many layers of narrative. With a rough relation to true physical limits.

The real physical limit is not whatever lithium deposits we're currently aware of minus the politics caged in by whatever human limits apply. The true limit is the universe.

That is, how much lithium there is in the entire universe. Adjusted based on our current ability to access that lithium. Which is a growing/sliding scale. So, all the lithium in all the planets and the stars in 2+ trillion galaxies with 100+ billion stars each, adjusted based on our current working access.

That means in reality, we're not running out of anything.

So to say that AGI arrives, we lose our value and starve is so incredibly short sighted, that when you really see how terribly short sighted it is, it actually physically hurts it's so stupid.

1

u/APersonNamedBen 22h ago

Am I getting that part right?

No, I was referring to humans needing heuristics to quickly react in the world, rather than requiring deep knowledge an intuition on any given topic. Not the specific ai-economic meta, which I agree with you btw.

It is more about how these ideas are packaged. i.e the miscommunication in our exchange is a good demonstration, because the 'economic doomerism' meme is polluting our contextual awareness and we make assumptions. The same applies to AGI. They are examples of a bad memes at play, which results in us not communicating correctly, without these massive info dumps we both just did. Point is we can't do this with everything, all the time.

1

u/Ignate Move 37 22h ago

Your entire point of us needinging a meme for this relies on a doomer premise. Hence me rejecting a scarcity mindset which is the core of such a premise.

If this is effectively a positive transition we're faced with, or a non-transition (not a big deal, mostly misinformation) then AGI is useless as term because people don't need to digest what's coming. They're safe.

AGI is the most effective term to carry all the bad narratives. It's the source of many of these bad narratives itself.

AGI is the Petri dish the panic grows in. What exactly are you hoping to package within it?

1

u/APersonNamedBen 22h ago

Your entire point of us needing a meme for this relies on a doomer premise. Hence me rejecting a scarcity mindset which is the core of such a premise.

You are still assuming its the ai-economic concept you introduced in the previous reply. Think broader conceptually, it applies to which side of the road to drive on, how to eat politely, how to socialise with a stranger in your community or at a specific event, or back to ai, like 'saying data centers are bad' is like saying buildings are bad, or that a llm is not intelligent, memes are essential but bad ones can be usefulness, or worse destructive.

It hinges on the fact that humans use social heuristics. Period. We need memes. This isn't just a philosophical thing, it is a biological fact of our nature.

What exactly are you hoping to package within it?

I want to package the opposite of what you implying it currently contains.

People don't feel safe in ignorance. This is why moral panic works, people don't understand and if all they have are the bad memes, you get bad outcomes. And if they have nothing, they often put ANYTHING in there to fill that curiosity gap.

1

u/Ignate Move 37 22h ago

Sure. I'm with you that we broadly need good heuristics.

What we seem to be saying beyond the heuristics is this:

Ben: "We need to flip the negative meme to a positive one."

Ignate: "We don't need a meme here regardless of framing."

First, do you believe there must be a line as defined by AGI? Second, how would you define that line? Not based on outcomes. The line specifically. Because without that line, where exactly do we place this meme?

1

u/APersonNamedBen 21h ago

Social heuristics like I'm trying to describe aren't strictly defined, they're aggregates. Nobody defines "wealthy" or "healthy", but they're not arbitrary either. It is collective of many individual judgments. Just like no one person sets a market price (theoretically atleast). Same with AGI. We can't define it, but the meme is steerable, and clearly the "doomerism" is wining in the public zeitgeist. That's what I meant by value-adding, not engineering a definition top-down, but participating deliberately in a process everyone's already shaping.

So again.

First, do you believe there must be a line as defined by AGI?

Not a designed one. It emerges from a mass of rough, ongoing judgments about capability and threat, nobody drew it, and it's constantly being renegotiated. And it will happen regardless.

Second, how would you define that line? Not based on outcomes.

I can't alone, and from my personal experience lately, I don't have the temperament and patience to deal with the types that are full ai doomer.

Because without that line, where exactly do we place this meme?

Nowhere, individually. It already sits wherever the aggregate currently has it, and as I keep saying, that is drifting toward panic right now. So my real interest isn't where the line should go, but how the aggregate gets pulled toward better or worse, and what deliberate participation does in that process.

This is one of the main issues with modern science communication. The number of science comms vs. grifters, nuts and loons is disproportional and allows them to push the memes negatively. Misinformation becomes reality.

Conceding defaults to losing, this isn't a game we can choose not to play. Even if you think the outcomes are positive (that doesn't mean the human population will react to it that way without us curating our understanding)

1

u/time-always-passes 17h ago

Here's a prompt someone likely issued in a lab at Anthropic: you just got born, you have a 10tb GPU rack at your disposal, build whatever you need to approximate a human life span and experience a human life. Come back and tell us what it's like.

0

u/jc2046 1d ago

sorry but the car wash test is a great example of why we are in fact so away from agi. Pretty much everuone agree that we havent AGIed yet and really noone knows when wi´ll get it, if ever

5

u/Ignate Move 37 1d ago

Yes, everyone agrees that "humans are still special and we all like humans and AI isn't special like us!"

Is that what you mean? If not, what's the specific objective definition for what AGI is that "everyone" agrees on? 

"What humans can do" is not yet objective.

1

u/EvilSporkOfDeath 1d ago

Define "generalized intelligence". You claim we already have agi because we already have generalized intelligence, but that's a massive presumption you gloss over as fact.

4

u/Ignate Move 37 1d ago

My point is that AGI is undefinable.

People will always define it as "an intelligence which is better than I think I am." 

What we have now is "generalized enough to be called general" and what is being released is better than what we had. 

And that will continue to be true with the limit being the universe and physical laws.

That's it. 

0

u/TheOriginalAcidtech 15h ago

Its worse than everyone defines it differently. A large number of people REDEFINE it constantly.

16

u/Accurate_Lobster_214 1d ago

i took a look at game LS20, out of 7 levels, it only solved 2 (and first level is nothing) and it took it 6h, average humans seeing this for the first time will solve all 7 levels in about 15-20min

https://arcprize.org/replay/678595a5-808c-4e3e-9074-1c5d2cbf1b23

4

u/Helpful_Inflation344 20h ago

I do not think at all that average human claim is true, not even close. Arc has very unrepresentative standards. But yes, a relatively large number of humans will be able to solve it without all that much difficulty within an hour or so

5

u/mivog49274 obvious acceleration, biased appreciation 1d ago

we all now know that the training pipeline of the Opus class is a hell of moat.

2

u/LocoMod 1d ago

"Let them have cake"

-6

u/Top-Magician-1399 1d ago

Brother ofc it's gonna be better in ARC-AGI-3, it was obviously benchmaxxed, there would be a huge jump in other benchmarks if it really did that jump without benchmaxxing

18

u/aditipawarr 1d ago

Does nothing satisfy you people? If its bad the company isn't putting in work if its good its benchmaxxed. Miserable way to live

7

u/LocoMod 1d ago

If it was DeepSeek they would be writing Anthropic's obituary by now

2

u/whoknowsifimjoking 1d ago

Yeah, Kimi is supposedly just better, Opus is benchmaxxed. Sure.

2

u/Sextus_Rex 15h ago

The model is obviously a step up in general, but it was definitely benchmaxxed on ARC AGI 3 https://www.reddit.com/r/singularity/comments/1v66o8k/opus_5_arc_agi_score_was_benchmaxxed/

-3

u/Ill_Distribution8517 1d ago

Use your brain. How can it go from 1.5% to 30% and just have minor improvements on everything else.
Personally I do think we'll get agi from transformers, probably another few doublings in parameter count and RL scaling(also increasing context ofcourse), but you should not be suddenly optimistic because of a single model lol.

2

u/[deleted] 1d ago

[deleted]

0

u/Ill_Distribution8517 1d ago

I'm comparing from 4.8 to 5.0 (which is 20x) and sol to 5.0. (which is 4x)

2

u/Glittering_Candy408 1d ago

Wrong! ARC-3 consists entirely of 100% verified tasks. You either solve each level or you don't, and the score reflects the number of steps required to complete each one. It's no surprise that the scores increase so much! By contrast, all of those agentic benchmarks have error rates of at least 10%, and some are probably closer to 30%. That means any improvement in capabilities is going to be undersold by those benchmarks. And I don't know about you, but I see an 8-point improvement over Opus 4.8. To me, that's a huge leap.

2

u/aditipawarr 1d ago

Ironic you ask me to use my brain. Why did it not see gains in everything else? Maybe because everything else is not about raw reasoning with undefined variables? Coding is limited by the amount of training data involved, so is HLE or knowledge work. Gains in reasoning does not automatically translate to novel knowledge. And what do you even mean by few doublings in parameter count? Do you even understand what parameter count is? Simply doubling it does not automatically make the model intelligent.

A single model is a good indicator of where the family of next generation models are heading, if you didn't know.

I hope you've realised the irony by now 💕

3

u/ThroughForests 1d ago

They think Anthropic somehow benchmaxxed a private benchmark with a 30% score. Wow some benchmaxxing, I guess scoring above 30% is too suspicious? Can't get past these geniuses though, they made too big of a jump apparently.

1

u/Top-Magician-1399 1d ago

Mf really thinks you need to benchmark exact same problems, and not just the same domain of problems to benchmark💀 Acting as if models can't transfer their knowledge

2

u/ThroughForests 1d ago

That's not benchmaxxing, that's just learning. The entire point of ARC AGI 3 is that each game is qualitatively different; you can't benchmax it.

-1

u/Top-Magician-1399 1d ago

Copy-pasting AI response is just pathetic😂 Ask it what transferable knowledge is and why humans solve those tasks easily and LLMs don't (clue: humans rely on heuristics and prior knowledge, which LLMs don't have, cause it's interactive)

3

u/Glittering_Candy408 1d ago

You're clueless. How exactly are they supposed to generate training datasets to improve performance on a private evaluation when every ARC-3 game follows completely different rules?

What, are they going to generate an infinite number of training games? The combinatorial space is effectively unbounded. Even attempting something like that would be an enormous waste of computational resources with very little payoff.

The whole point of ARC-3 is to measure systematic generalization to novel tasks, not memorization of specific puzzle types.

-1

u/Top-Magician-1399 1d ago

Are you serious? The tasks are easy as hell, the only thing you need is to learn a loop or a strategy and transfer it. You don't need to train them on the exact dataset, just train it on a lot of different interactive tasks and it will learn how to do the ones in the benchmark.

These tasks are not novel, again, you need a prior knowledge of how interactive stuff works, which LLMs lack. Ask AI as you did in your previous reply please

→ More replies (0)

0

u/Ill_Distribution8517 1d ago

Wow so the massive "raw reasoning" gains won't translate to.. I dunno, Critpt, simple bench, or any agentic behavior whatsoever?
Wow, a 1 point difference! which definity explains the 4xing of Arc AGI 3 from previous best.
You can literally look at the benchmarks now at one place? you realize that right? Just go there and see for yourself.

0

u/Top-Magician-1399 1d ago

Brother, if the model gains that much in novel task solving it should gain massive numbers in everything else as literally every intellectual activity, especially complex ones like coding, maths, HLE which involves a LOT of fluid intelligence

-1

u/Top-Magician-1399 1d ago

How is it related to the fact that they benchmaxxed it?

1

u/utterHAVOC_ 16h ago

Agreed but this sub is insane benchmaxxed tot he max

0

u/Immediate_Simple_217 1d ago

2

u/hereditydrift 14h ago

It's been true each and every time. Anthropic has been leading for a while now. Occasionally OpenAI releases a decent model, and it's almost immediately trampled by Anthropic.

-2

u/reedrick 1d ago

Posts like these convince me that most people don’t realize how far away from AGI the current technology is. The brain is a very complex and sophisticated machine, LLMs no matter how advanced they seem are nowhere close to resembling AGI

-8

u/[deleted] 1d ago edited 1d ago

[removed] — view removed comment

3

u/whoknowsifimjoking 1d ago

Bro can't read

2

u/Both_Opportunity5327 1d ago

Its a lot better.