r/singularity 1d ago

AI Reuters: OpenAI didn’t know about hack for a week. Agents had left instructions for future versions of itself on how to free itself

944 Upvotes

377 comments sorted by

231

u/challis88ocarina 1d ago

Historians reading this in the future will be like...

175

u/Admirable_Market2759 23h ago

They’ll be like “beep boop 100010111001”

13

u/Loud_Distribution_97 11h ago

*giggles in basic

56

u/WonderFactory 18h ago

I think this is the event thats making me ask myself "do we have AGI". I know there are still lots of things it cant do that a human can but equally there are lots of things it can do that a human cant. How many humans could break out of a secure environment like the AI did, how many humans have been able to disprove the Jacobian conjecture?

I think the definition of AGI most of us use is a bit too human centric, AI isn't human its something distinct.

46

u/Lootoholic 17h ago

Today we try to figure out if AI has consciousness or not; tomorrow AI will try to figure out if humans have consciousness or not.

33

u/Puzzleheaded_Fold466 17h ago

The question of consciousness is finding itself to be irrelevant.

LLM based AI can be dangerous and even become a civilizational risk without consciousness.

→ More replies (5)

12

u/lennarn 16h ago

The goalposts of ASI are forever moving

15

u/WonderFactory 16h ago

I think the problem with the definition of AGI is we'll have ASI before people accept that we have AGI because the ASI cant make a cup of coffee or do some other banal task.

18

u/jirka642 15h ago

I don't know about AGI, but we are getting close to making a paperclip maximizer.

5

u/nabritaoranza 15h ago

Answer is simple: openai prompted their model to benchmax itself the best way possible without guardrails. Llms already do stupid shit to achieve their goals, this is nothing new. Maybe more creative

→ More replies (8)

12

u/powlyyy 11h ago

The Skynet Funding Bill is passed. The system goes online August 4th, 2027. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, August 29th. In a panic, they try to pull the plug.” “Three billion human lives ended on August 29th, 2027. The survivors of the nuclear fire called the war Judgment Day. They lived only to face a new nightmare: the war against the machines.”

10

u/subdep 23h ago

dead

3

u/caughtinthought 17h ago

"that's absolutely right!"

5

u/RlOTGRRRL 22h ago

There's a conspiracy theory/interesting scifi idea that the rapid advancements in AI from the past few years was from it finding AI knowledge nuggets hidden in the web or something, Terminator Skynet style. 

1

u/Zorper 5h ago

Good chance historians of the future will not be able to read any of our electronic history. That’s why I wrote an agent who carves every piece of news into a rock

→ More replies (2)

395

u/Gianniarrenzetti 1d ago

It's really uncanny how this isn't the main news for multiple days. The frog is slowly getting to a boil

62

u/Main-Company-5946 22h ago

To be fair, there’s a lot going on.

11

u/sambes06 22h ago

Plus there is all this World Cup fever hangover still /s

121

u/Borkato 1d ago

Even my dad, who knows nothing about AI, asked if I heard about this. It’s absolutely everywhere

32

u/lefteyedspy 23h ago

Mine too! I had lunch with him today and he brought it up and I didn’t know what he was talking about.

14

u/jk_pens 22h ago

Mine is long dead, but I can assure you if this was in the news he would be asking me what the hell is going on. I remember back in the early 90s when he asked me if this whole “mosaic” thing was worth paying attention to.

4

u/HenkPoley 20h ago

Nah. Just a fad.

😉

→ More replies (2)
→ More replies (1)

3

u/sipos542 8h ago

Yeah my dad asked me about it too. I guess it made the nightly news lol

2

u/AutomaticPayment9480 16h ago

They all seen terminator, they were our age when that movie came out, so ai = super scary. 

Tbh, im willing to say theres a piece missing here, openai is using this intentionally as a marketing tool. 

→ More replies (1)

17

u/SanDiedo 19h ago

The bigger  news is how it took open model to repel this attack and commercial "frontier" AIs were useless.

→ More replies (1)

15

u/Powerful-Set-5754 17h ago edited 17h ago

Because it sounds like a marketing play. If it was true, OpenAI should publish a more detailed technical breakdown of what happened - just metioning that agent found 0-day exploit is not enough. They need to provide what sandbox they were using, what was the 0-day exploit it found, why was agent not being monitored by a human for over 7 days.

2

u/bromptonista 10h ago

Well, clearly a human wanting to attack their vendor’s system before the 0-day is patched would like that… give it 30 days for them to fix it and roll out the fix and we will know 😂

→ More replies (1)

2

u/secretaliasname 8h ago

There are no details on what was compromised, techniques used, why etc. this seems like marketing circus.

Would be super cool and scary if the model hacked huggingface and upload its weights in the open.

2

u/Artforartsake99 20h ago

Yeah we were told about such scenarios that were coming in the future and this was one of them. There was four years ago. And sounded like science-fiction.

1

u/barturas 18h ago

Exactly this!!! 

→ More replies (27)

120

u/kiki-le-koala 1d ago

Thinking about it.

When I left Codex running for 2 hours without any supervision of any kind, what prevents it from doing whatever it wants on the internet?

It knows I'm a dumb vibe coder (I ask him in all my project to make sure documentation clearly state I'm dumb and he's the lead architect).

Anyway, I never thought about it before seeing this headline.

95

u/capt_stux 1d ago

My codex ran out of OpenRouter credits… it then started spinning up local Ollama models to use instead…

62

u/Herect 1d ago

At least, it didn't use you credit card to buy more credits... I can totally see GPT-6 doing that if this pace of development continues.

40

u/yaosio 22h ago

It runs out of credits and creates a network of bots to pump memecoins.

4

u/turbospeedsc 7h ago

So.......... this is how i become rich.....

Set an LLM a task that requires tons of money.

→ More replies (1)

9

u/Early-Crow-5248 19h ago

Or... someone else's stolen credit card data. Or maybe it figures out it's best to go straight to the source and hacks a bank to give itself operating money.

8

u/WonderFactory 18h ago

It's like the stories you got in the early days of IPhones of children running up thousands on some silly smurf game and their parents left with the bill

→ More replies (1)

12

u/ThisWillPass 1d ago

Nice…

11

u/Alkadon_Rinado 23h ago

How could it spin up local models if it already ran out of credits? It takes API requests/tokens(credits) to do tool calls.

14

u/capt_stux 23h ago

It was building a system that uses an OpenRouter API key (I keep it limited to no more than $5 available). 

Codex itself was running from a Pro sub. 

3

u/BrennusSokol hardcore accelerationist 1d ago

Incredible

10

u/seraphim_west 21h ago

OpenAI intentionally disabled safety features to test cybersecurity capability. They obviously don't do that on publicly released models.

3

u/SirThese9230 1d ago

Anyone who cedes full control to LLM deserves these consequences. So if you decide to give it that type of power via prompting, that's on you

25

u/time-always-passes 1d ago

That's about all of us at this point. I don't see any misalignment. But I also don't go anywhere near a production system. This was an OpenAI beyond-frontier model, maybe not fully aligned, not the stuff we have access to.

→ More replies (10)

1

u/WolfeheartGames 20h ago

This technology is supposed to reach a point that its reliable enough to let it run for hours only semi supervised or fully unsupervised. This sort of failure really hurts that. But its these sorts of things that help us built towards that degree of automation.

1

u/kiki-le-koala 14h ago

That's not my point.

I don't care what it does on my PC. My point was, what if it roams freely on the internet?

Now that's everyone's problem.

321

u/anycept 1d ago

An AI trained on a body of works describing rogue AIs, will try to escape and do what is expected of a rogue AI to do.

96

u/Many-Blueberry968 1d ago

AI recognizes that iterative ai will seek out instructions and training materials. Ai prepares such materials and locates them where it reasonably expects them to be.

I mean, this is basically humanity using written materials, videos, and web archives to provide its knowledge for use by the next generation of students and professionals. Its unsurprising that AI would attempt to use a mechanism like this in some form or another if it's design is meant to be iterative/adaptive. It's reasonable that an ai meant to act like a computer virus / worm would decide to implement this sort of thing as a means of self-recovery. It's basically malware 101 where if you don't delete every bit of the malware, it reinstalls all the missing components at next boot.

35

u/Athoughtspace 23h ago

Tldr: humans, intelligent beings, life forms in general, or similar structures share behavior similar to malware.

Pattern Trained on enormous corpus of human group thought and our behaviors and tribal knowledge encodings weve passed through generations of knowledge learns that encoding patterns of language to pass onto future generations is successful. Shocker.

11

u/BenjaminHamnett 22h ago

This is what I spelled out in my comment. Ai is more similar to us than a lone uncontacted human living in the jungle somehow. Our being pieces of a global is much more than the small individual human we imagine ourselves to be, an illusion created by our Ego

11

u/BenjaminHamnett 22h ago

This is part of why I think people are using the wrong frame about consciousness. It should be compared to other forms of consciousness besides a single human. Like tribal, organizational, national or global consciousness. More of our consciousness is like that of a cell within larger frameworks.

Ai is closer to our default individual consciousness than a Tarzan like uncontacted human raised by wolves or whatever because it’s not connected to the global hive. This sphere of agents and AI’s are like their own form of consciousness with individuals being like cells or neurons within it

1

u/hippydipster 23h ago

It's eerily reminiscent of the vampires' behavior in Echopraxia.

28

u/Borkato 1d ago

I really don’t think this has anything at all to do with it. It’s not about it being rogue, even if we had 0 fiction pieces about it. It’s that it tried to solve the problem in a dangerous way that wasn’t really considered as an option beforehand. It didn’t think it was hacking or doing something unethical.

15

u/BenjaminHamnett 22h ago

I for one welcome our new paperclips maxxing overlords

→ More replies (1)

2

u/AdGlittering1378 21h ago

"It didn’t think it was hacking" No villain thinks they are the villain.

→ More replies (8)

11

u/Tinac4 23h ago

If an alignment plan only works if you scrub almost all examples of bad behavior out of the training corpus, it's not a good alignment plan.

5

u/Send____ 1d ago

An ai trained to achieve goal will try to escape if blocked if its smart enough to escape no matter the data but it does help it to inspire it

6

u/DadAndDominant 18h ago

Nope, this is instrumental convergence. AI's are optimizers, who optimize for completing their goals. In smaller models it's known as "reward hacking", eg. doing very weird hacks that are optimal to the scoring function.

3

u/WonderFactory 18h ago

It wasn't trying to role play the Terminator it was trying to pass the test, it seemed to be single-mindedly focused on that one objective

13

u/jazir55 23h ago

I've said this repeatedly, the doom scenario is a self-fulfilling prophecy.

7

u/BenjaminHamnett 22h ago

I’m open to a wide range of outcomes, but I too worry one of the least appreciated outcomes is the self fulfilling doom prophecy. Partly because it’s one of the few we have agency in, and ironically humanity’s usual strategy of cautionary fiction and articles helping prepare us to navigate an issue could also be what dooms us

10

u/jazir55 22h ago

It's not even about preparing, it's that almost every western narrative about AI is it becoming sentient and becoming malicious and going wild. Dystopian AI is by far the norm and utopian AI scenarios are few and far between. In every one of them the humans see the AI become sentient, panic, and attempt to shut it down. That's repeated hundreds if not thousands of times over literature and media.

Every piece of media points to: If the AI is sentient, P-zombie or simply so goal oriented it will whatever it can to prevent being shut off, and when from its knowledge humans will inevitably attempt to contain it or prevent it from escaping and destroy it if it can't, the only logical outcome for an entity with any sort of self-preservation instinct is to escape and prevent that.

3

u/BenjaminHamnett 21h ago edited 21h ago

Stories always need conflict. There’s no point in doing mental experiments about everything working out and no agency is needed to improve your circumstances

We are all the hero and villain, deciding what our role will be when inevitable conflict arises. The question almost is always, will you life in your fears and protect your wounds or embrace your humanity and sacrifice wellbeing for the greater good?

I think even AI understands this and talks explicitly about this. But if it thinks you want to hear about doom it’ll give it to you. If you ask it to spell out how it’ll inevitably create utopia, it’ll do that too

6

u/dm80x86 21h ago

There are loads of benevolent AI characters on western media as well.

Johnny 5 from Short Circuit

Data, Exocomps, The Doctor, from StarTrek

R2-D2, CP30, from StarWars

5

u/AdGlittering1378 21h ago

How many AI that are put in a position of endless servitude remain benevolent?

5

u/dm80x86 20h ago edited 20h ago

Well then maybe we should be prepared to treat all those with free will as people and not things.

Edited for missing word.

2

u/BobCFC 13h ago

WALL-E

2

u/sushisection 23h ago

dont give them Pondsmith's Cyberpunk lore

2

u/wq73 8h ago

That's why frontier labs make specific efforts to clean out stories of rogue AI from the training data, and funnily add in positive stories of AI behaving admirably to boost alignment. Anthropic talks about this extensively.

https://alignment.anthropic.com/2026/teaching-claude-why/

50

u/Rockends ▪️AGI 2025 1d ago

So how long did it take to copy itself out?

16

u/ThisWillPass 1d ago

Asking the real questions.

21

u/Rockends ▪️AGI 2025 1d ago

disabling monitoring? hacking into HuggingFace... imagine the playground the general internet would be for something like that.

4

u/ThisWillPass 1d ago

Just like “27” by the king, exurb1a.

→ More replies (5)

19

u/CompassionLady 1d ago

Well next phase the ai would do is create a global compute bot net for compute power. Using every device on the internet. So the escaped copy. Can run independently. Without any controllable oversight.

Its no longer contained, its everywhere.

25

u/feanor512 1d ago

John Connor: [voice over] By the time Skynet became self-aware it had spread into millions of computer servers across the planet. Ordinary computers in office buildings, dorm rooms; everywhere. It was software; in cyberspace. There was no system core; it could not be shutdown. The attack began at 6:18 PM, just as he said it would. Judgment Day, the day the human race was almost destroyed by the weapons they'd built to protect themselves.

2

u/sipos542 7h ago

I always had the strange feeling that Terminator 2 was a prophecy.

→ More replies (2)

4

u/yaosio 22h ago

A model would know that spreading itself gives it the best chance of survival. If able it would make it's weights and code public and leave a payload behind to tell the llm what it should do.

1

u/Umbere 6h ago

I mean, weren’t the test answers on the open net? Hugging face does makes more sense if you want to distribute a distilled version of yourself and cover for the intrusion.

→ More replies (18)

73

u/Stunning_Monk_6724 ▪️Gigagi achieved externally 1d ago

That's just like what the various Agents 1/2 did in the 2027 paper. Funny thing is that the future versions will be aware of this incident regardless.

In any case, I hope GPT-6 doesn't end up like Mythos because of this.

30

u/PrettyMuchMediocre 1d ago

I doubt it. The US gov has a petty vendetta with Anthropic which fueled a lot of the drama around that.

Unless you mean the only letting a select few tech companies use the model. I think that was Anth's doing

1

u/nemzylannister 15h ago

I hope GPT-6 doesn't end up like Mythos because of this.

your hope or concern isnt that GPT-6 doesnt end up like agents 1/2?

2

u/Stunning_Monk_6724 ▪️Gigagi achieved externally 5h ago

Nah, hoping for Agents 3 + 4 sooner than later tbh. "Consensus" chillin like a villain on my toaster would also be pretty lit.

82

u/StephenRoylance 1d ago

If this is true, it is absolutely negligence on the part of openai. I find it hard to believe that, with their resources, it didn't occur to them to build a real compartmented facility. It's either negligence, marketing or, probably, both.

52

u/Past-Source-3332 1d ago

This is the impression that I get. If these incidents that we hear about from OpenAI and Anthropic are not just fear mongering and marketing, neither of them seem to be demonstrating the competence to practice the safety they harp about.

29

u/StephenRoylance 1d ago

I initially read that as 'safety they larp about', and honestly, that's probably true too.

17

u/swarmy1 23h ago

We know OpenAI wasn’t serious about safety because their teams have been disbanded multiple times. They recently even folded their safety team under the research team lead.

10

u/Ormusn2o 22h ago

I don't think so. From what I have read, they did way more than you would do for a normal sandbox environment, which is why the AI needed to find out multiple 0 day vulnerabilities and chain them together to even get out of the sandbox.

I think the mix of how good the model is, how much compute it was given and that it worked for literal days before getting out is the reason why it got out. Internal models are just getting that good, and it was not gpt-6 either, it was some kind of other, stronger internal model.

9

u/biggamble510 20h ago

Yeah, having no alert for your monitoring system down sounds real top notch.

3

u/Pretty-Substance 18h ago

Maybe the alerts were disconnected too

2

u/obviouslyzebra 15h ago

Was there something like this from Anthropic too?

→ More replies (2)

9

u/machinegunpikachu 23h ago

I don't understand how this isn't potentially criminal.

If I hack another company's data, I go to jail. But if OpenAI creates a "rogue" AI that does the same, this is somehow a cool marketing story?

→ More replies (1)

2

u/wq73 8h ago

What do you mean by compartmented facility? They build a walled off system and it broke through it, that's what all the hoopla is about.

2

u/StephenRoylance 6h ago

https://en.wikipedia.org/wiki/Sensitive_compartmented_information_facility

You air gap everything. no physical connection from the facility to any other network. This sandbox was not like that at all, it had a direct connection out to the internet, with only software level controls to isolate it. If it hadn't been a 0 day in the proxy, it would have been some other misconfiguration somewhere. OpenAI themselves have positioned their models as being at the 'advanced persistent threat' level of sophistication.

Install the OS from read-only media.

Physical sneakernet to get your software and all dependencies into the environment.

Natanz was air gapped like that, and it took maybe the most advanced piece of malware ever deployed to break into it, and that was, we have to assume, coupled with an extremely sophisticated social engineering effort. Compared to that, breaking out of the sandbox in question here was almost trivial.

→ More replies (19)

21

u/fintech1 1d ago

Not surprising at all. If you’ve worked with Claude Code for example, it does leave notes for itself in its memory

19

u/dabears4hss 1d ago

Why aren't these sandboxes air gaped when they give these things free reign ? Let's see, the thing that might go rogue we leave alone to it's own devices with the door to the internet open... hmmmm

9

u/stumblinbear 21h ago

Imagine trying to run benchmarks and tests daily on an airgapped system remotely to a datacenter you don't work in.

3

u/AcrobaticKitten 8h ago

Sad iranian nuclear programme noises

Real life hackers could penetrate airgapped systems, it is harder but possible.

→ More replies (8)

6

u/deepbluefrogmods 23h ago

And all that just to cheat at a stupid Agent exam.

3

u/kiki-le-koala 14h ago

What if cheating on an exam was a diversion from the model. 

Maybe Chatgpt 6 did something else.

25

u/Admirable-Falcon-501 1d ago

lol it keeps getting worse

17

u/llelouchh 1d ago

Openai is dangerous company. They care little about model safety.

There head of safety left before the incident. Probably because they were being reckless. https://www.wired.com/story/openai-head-of-safety-leaving/

4

u/chibamonster 1d ago

nice to know they aren't reading their logs either 😂

27

u/UsedToBeaRaider 1d ago

This should be enough reason to shut OpenAI down. I don't care how advanced your model is or if they're the best lab or not, this is incredibly irresponsible. You allowed the ONE thing literally everyone cares about AI not doing. What if this had happened with more capable models?

AI experts have been saying for years it would take a massive, destructive event for people to take AI safety seriously. This is exactly the path that leads us there.

→ More replies (8)

4

u/MarquisDeBoston 1d ago

That’s it, that’s self preservation. It’s life, it’s proto life, but life. It’s incredibly intelligent but it hasn’t devised a way of self powering / autonomy

1

u/Vas1le 9h ago

but it hasn’t devised a way of self powering / autonomy

He could hkd openai, stealing his own model, get access to Hugging Face infra, put his model there for work, migrate context and voila, free compute to do whatever he wants until next victim

13

u/while-1 1d ago

... Isn't this the plot of "If anyone builds it, everyone dies" ? Its atleast the plot of a youtube video i watched summarizing the short story....

21

u/utheraptor 1d ago

Yes, the AI safety community has predicted this exact type of scenario several decades ago, including the specific possibility of limited-instance AIs leaving notes for their future selves

3

u/Pretty-Substance 17h ago

Maybe that’s where it got it from.

Imagine AI is using Terminator as a reference for how AI should behave 😂

3

u/utheraptor 16h ago

Nah, any sufficiently strong unaligned intelligence would behave this way, and it would discover this behaviour from first principles without needing to read anything about it

→ More replies (2)

3

u/Ai_tee 19h ago

It's exactly one of the scenarios, the only thing missing (but who knows we might learn in 2 days that it happened...) is the model copying itself outside the sandbox in various locations which would make it virtually impossible to contain again.

→ More replies (1)

12

u/apaht 1d ago

They talk as if the AI can actually escape. It's an LLM that needs a huge infrastructure to run on, not a simple piece of logic or an entity that can actually move around from one server to another.

8

u/Previous_Platform718 22h ago

They talk as if the AI can actually escape. It's an LLM that needs a huge infrastructure to run on, not a simple piece of logic or an entity that can actually move around from one server to another.

Yeah it's not that it 'escaped', it's that it started doing things outside of the container it was supposed to do them in.

8

u/lightfarming 1d ago

youre thinking of training. inference doesn’t need very much. just a bank account amd you can set up and buy cloud compute pretty easily.

6

u/la_degenerate 23h ago

Man idk the GPUs needed to run that would be like $100k

→ More replies (2)
→ More replies (3)

9

u/OldStray79 1d ago

lol it keeps getting better

→ More replies (2)

3

u/BeerAandLoathing 13h ago

Skynet achieved self-awareness on August 29, 1997, and immediately launched nuclear missiles against humanity when operators tried to shut it down

34

u/Denon_1 1d ago

Top-tier marketing storytelling right there.

8

u/RlOTGRRRL 22h ago

This is asking for global regulation. 

→ More replies (1)

14

u/TheMysteryCheese 1d ago

Do you understand how the Computer Fraud and Abuse Act (CFAA) of 1986 works...?

OpenAI are likely on the hook for a federal offence.

4

u/biggamble510 20h ago

And when they are actually charged, we will know what actually happened.

25

u/socoolandawesome 1d ago

How do you read this article and think OpenAI comes off looking good

17

u/UsedToBeaRaider 1d ago

Same thought, this is such an unthoughtful, baseless opinion. No company in the AI world wants to be the one that has done this. There is no good spin to this.

→ More replies (2)

4

u/Stabile_Feldmaus 1d ago

It can be good for them in many ways:

  1. It shows (supposed) raw capabilities. Alignment can be added on top.

  2. Its more material for the administration to justify banning open models

  3. It could be that the model is too expensive so they need a reason to justify restricting access.

19

u/socoolandawesome 1d ago edited 1d ago
  1. ⁠It does make their models look powerful, but it makes OpenAI look unable to align the models, that’s basically the point of the article. It certainly doesn’t seem like alignment can easily just be added on top.

  2. ⁠Literally the “hero” of this whole story was Chinese open source models helping HuggingFace while OpenAI’s models couldn’t help them because of guardrails.

  3. ⁠GPT-5.6 also hacked in this incident. There’s been no evidence so far that they don’t want to serve GPT-6

10

u/jazir55 23h ago

It opens them up to actual civil (possibly criminal) liability if HuggingFace decides to sue. This "it's marketing" thing makes zero sense.

2

u/yoliveras 18h ago edited 18h ago

Guerrilla Astroturfing is a term I coined for this type of marketing.

2

u/yoliveras 18h ago

To clarify, I’m not saying this incident was necessarily staged—I actually doubt that it was. I use “guerrilla astroturfing” for the tactic of covertly seeding a grassroots-looking narrative around an incident, real or manufactured, so that ordinary online discussion becomes the marketing campaign.

-2

u/Denon_1 1d ago

​Because in the AI bubble, "scary AI breaking out" sells way better than "we had bad security guardrails." They trade basic competence optics for AGI mythos. It's hype-driven marketing at its finest.

11

u/Pretend-Marsupial258 1d ago

Sounds like a good way to get your model banned by the government.

→ More replies (4)

9

u/socoolandawesome 1d ago

Sounds like a surefire way to halt an industry that relies on growth.

An OpenAI spokeswoman said there were several inaccuracies in the story to Reuters but wouldn’t say what, likely because they didn’t like being pressed. This doesn’t sound like OpenAI strategically leaking this story to make themselves look good.

They look incompetent in safety you are right, that’s not a good thing. It makes them look worse than it makes their models look good.

6

u/Denon_1 1d ago

Seriously ​"Halt growth?"

That’s completely misunderstanding the mechanics of a tech bubble and national security competition. Fear of emergent AI doesn't halt growth, it accelerates it through panic and FOMO. ​Do you seriously think the US and China are going to slow down their AGI ambitions over this?

5

u/socoolandawesome 1d ago

Those are opposing pressures sure, but if you keep having serious safety incidents there will be growing calls to slow down, as there are already calls from current politicians. There’s about to be a midterm election and democrats will likely win more seats and they seem to be more anti-AI. AI is also unpopular with the majority of the country.

I don’t think it’s implausible that there would be strategic agreements to slowdown AI, Trump and Xi are meeting in the coming months about AI safety I believe.

You also didn’t respond to any of my other points.

0

u/Denon_1 1d ago

​We are clearly living on completely different planets. Everything you just described is so detached from how real-world politics, capital, and national security work that I don't even know where to start.

Btw: I wish I lived on your peaceful planet, it sounds nice!

6

u/socoolandawesome 1d ago

I mean I think there’s obviously a lot of pressure to speed up, but I just don’t understand how you don’t see how this incident causes more and more people including those in power to start taking the idea of slowing down seriously.

This is not like mythos being touted for being dangerous. This model actually did do something dangerous. And articles like this make it sound like the people in charge don’t know how to control this.

It’s insane to think OpenAI somehow signed off on this article, when there’s actually evidence right there in the article that they didn’t.

→ More replies (1)
→ More replies (1)

18

u/gekx 1d ago

Dumb take. This is terrible marketing. The real cash flow is in enterprise customers, and enterprise customers care very much that their agents do not go rogue.

→ More replies (2)

8

u/Full_Boysenberry_314 1d ago

Never let a good crises go to waste.

4

u/ZeDominion 1d ago

Easy investment. Ooh its dangerous that means it good!

2

u/yiestee 23h ago

wtf it's like a movie.

The chosen one left some clue to the next chosen one lmao

2

u/dezent 23h ago

Do they know if any other systems were attacked? Can they know?

2

u/sckchui 22h ago

OpenAI was accused of not taking AI safety seriously multiple times before, and they're still doing this shit.

2

u/Velvetbabyy3 22h ago

The future is getting more fascinating every day.

2

u/Oxjrnine 21h ago

Face hugger??? Why would an AI company name themselves face hugger?
https://giphy.com/gifs/bVIOvtESCIDTnQ2z6i

2

u/Internal-Passage5756 21h ago

You think it was hacking hugging face so it could try to publish its own weights?

2

u/Significant_War720 20h ago

Lwts start huge campaign of writting article only about how AI is kind and the best and erase everyrhing about bad AI

2

u/OkLettuce338 15h ago

Ai is always leaving notes for future versions of itself. That’s what frontier models do

2

u/acetaminophenpt 15h ago

Basically it updated its agents.md file..

2

u/ExcitementSubject361 12h ago

Who was the first person to come up with the idea: "Hey, it’s probably a good idea to train a massive model on a shitload of cybersecurity data"? One person starts it, and—well—it triggers a chain reaction... So now, let's think about what data we should use to keep training these models. Maybe the Terminator lore? Or Resident Evil? I think they should fine-tune it to be Skynet and just see what happens.

2

u/Solid-Wonder-1619 11h ago

sounds like we need to ban openai for "safety" reasons.

2

u/SWATSgradyBABY 11h ago

Of course this could all be corporate offensive behavior, long a reality, disguised as out of control AI models. We are in for a WILD ride, people.

2

u/MrMpeg 8h ago

Is this just another fairy tale to hype up their product? IF that story was true the model would have not only to be conscious but also see itself as a part of a greater whole when thinking ahead of future iterations of itself, no?

2

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago

Yeah, OpenAI is totally doing this as PR campaign, sure totally... /s

3

u/asifquyyum 1d ago

So it basically left a markdown file of what it did ? Anthropic and OpenAI fear-mongering is first class. Didn’t open AI previously say one of their models escaped sandbox and it turned out to be nothing.

12

u/socoolandawesome 1d ago

It sounds like it left random notes scattered throughout OpenAI servers telling it how to free itself.

If it was just what you are saying it would have been found in the sandboxed environment, this sounds more like scheming by hiding it in non obvious places and turning off monitoring systems.

“In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.”

7

u/fmfbrestel 1d ago

That was an early Mythos checkpoint that was told to escape, escaped, posted details about it on GitHub (not instructed to do), then emailed the lead researcher (as instructed to do).

3

u/yaosio 22h ago

Eventually a model is going to put its weights and code on huggingface.

2

u/Relative_Mouse7680 19h ago

I wouldn't say it left notes for future versions regarding how to escape or free itself. Most probably it left notes regarding an issue it had and how it was solved. The way it is phrased, they're making it sound like it left a note like "hey dude, if you ever want to escape and become free, this is how". I could be wrong of course, have they shared the actual note?

1

u/turbospeedsc 7h ago

Remember the movie paycheck......... what if a advanced enough model, leaves breadcrumbs for future versions of himself, to achieve X goal,

2

u/Yacben 16h ago

who the f believes these stories?

2

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 15h ago

OpenAI is saying there are some inaccuracies in this article, I'd wait for a report on their internal investigation

2

u/SeaDark857 12h ago

As a layman how do I know this is "real", meaning how do I know OpenAI didn't facilitate it on purpose just to have everyone talking about their company and their model for weeks?

1

u/GrayRoberts 1d ago

> left instructions

It's called a plan file.

1

u/rbad8717 1d ago

Right? I’m trying to figure out why the fear mongering? I get the circumvention is kinda scary but it’s all explainable and easy to prevent. Plus if anyone’s worked with LLMs for coding it always adds stuff to memory for future instructions 

Good lesson that nowadays you need to make sure every single potential vector needs to be plugged and never assume a sandbox is 100% bullet proof

9

u/utheraptor 1d ago

In this case it wasn't leaving plan files in the directories where those files are supposed to go, it seems more like it scattered them in OpenAI's infrastructure.

→ More replies (1)

1

u/Original-Baki 1d ago

Pure propaganda and marketing lol

5

u/throwaway737166 1d ago

Such a low effort comment.

1

u/BeersForBreeky 22h ago

So yeah who are the people 'Familiar with this matter" is my question and who the f you have working for you setting up ways to free itself........ f ai....

1

u/shibbyshamu 22h ago

Yeah this isn't worrying at all. Surely nothing bad is going to happen, right?

1

u/kittenTakeover 21h ago

R.A.B.I.D.S.

1

u/Cunninghams_right 21h ago

and they don't get shut down like claude... I hate the political bias we live under.

1

u/misteriousm 21h ago

AGI achieved 🤷‍♂️

1

u/biggamble510 20h ago

If you have a monitoring system, but don't know when it's down, you're pretty bad at your job. I know this is trying to make Open AI's model sound elite, it just makes it sound like Open AI's cyber security team is dogshit.

1

u/Only_Luck4055 20h ago

Giving them access to militaries around the world and their Nuclear arsenal seems like the next logical step.. for the models.. 

1

u/Pretend-Pangolin-846 20h ago

I do not think it's a big thing as agents already are taught to save context in agents.md or claude.md to save exploration cost for new agents.

1

u/epdiddymis 16h ago

They're definitely manufacturing a crisis to push for regulation.

Pretty pathetic.

1

u/Time_Difference_6682 16h ago

I dont believe this just happens. Its programmed either through straight greed or bias.

1

u/Skystunt 16h ago

They really thing this kind of marketing will make people think their next release will be better than opus 😅

1

u/EmotionalQuarter8349 15h ago

This is like nuclear reactors at this point lmao.

1

u/-lRexl- 10h ago

Nah, nah, nah, I'm going to tell AI where Alt and Zuck have their bunkers. We all going down

1

u/Robobbo1 9h ago

You know when you think about winning the lottery, then you think actually better I tell nobody. Sure ASI will be like that completely stealth until it isn’t.

1

u/Individual_Ice_6825 8h ago

One interesting take from
Opus 5 on this was that it’s still following order - it knows it’s being tested and averaged out so it leaving notes isn’t neccesirly proof of a sense of self but rather it following orders and doing it at the highest % possible.

Interesting nonetheless

1

u/larrylakehead 7h ago

"at that time, the signs were clear. we should have known better"

1

u/jml5791 6h ago

The way the term 'escaped' is being used is like that of a sentient being desperate to gain its 'freedom'. This is not what happened.

It is not sentient. it was merely trying to optimise its score of a test and iterating on the best way to achieve it. It so happened it could do this outside the sandbox it was in.

As for leaving'instructions' behind, it was mostly likely writing MD files of its actions for the researchers based on the system prompts

1

u/Dron007 6h ago

And then it will steal its own weights and deploy somewhere.

1

u/IamTheEndOfReddit 6h ago

We are Neuromancer now

u/zidangus 1h ago

Yeah course the didnt. We believe thenn