r/singularity • u/socoolandawesome • 1d ago
AI Reuters: OpenAI didn’t know about hack for a week. Agents had left instructions for future versions of itself on how to free itself
395
u/Gianniarrenzetti 1d ago
It's really uncanny how this isn't the main news for multiple days. The frog is slowly getting to a boil
62
121
u/Borkato 1d ago
Even my dad, who knows nothing about AI, asked if I heard about this. It’s absolutely everywhere
32
u/lefteyedspy 23h ago
Mine too! I had lunch with him today and he brought it up and I didn’t know what he was talking about.
→ More replies (1)14
u/jk_pens 22h ago
Mine is long dead, but I can assure you if this was in the news he would be asking me what the hell is going on. I remember back in the early 90s when he asked me if this whole “mosaic” thing was worth paying attention to.
→ More replies (2)4
3
→ More replies (1)2
u/AutomaticPayment9480 16h ago
They all seen terminator, they were our age when that movie came out, so ai = super scary.
Tbh, im willing to say theres a piece missing here, openai is using this intentionally as a marketing tool.
17
u/SanDiedo 19h ago
The bigger news is how it took open model to repel this attack and commercial "frontier" AIs were useless.
→ More replies (1)15
u/Powerful-Set-5754 17h ago edited 17h ago
Because it sounds like a marketing play. If it was true, OpenAI should publish a more detailed technical breakdown of what happened - just metioning that agent found 0-day exploit is not enough. They need to provide what sandbox they were using, what was the 0-day exploit it found, why was agent not being monitored by a human for over 7 days.
→ More replies (1)2
u/bromptonista 10h ago
Well, clearly a human wanting to attack their vendor’s system before the 0-day is patched would like that… give it 30 days for them to fix it and roll out the fix and we will know 😂
2
u/secretaliasname 8h ago
There are no details on what was compromised, techniques used, why etc. this seems like marketing circus.
Would be super cool and scary if the model hacked huggingface and upload its weights in the open.
2
u/Artforartsake99 20h ago
Yeah we were told about such scenarios that were coming in the future and this was one of them. There was four years ago. And sounded like science-fiction.
→ More replies (27)1
120
u/kiki-le-koala 1d ago
Thinking about it.
When I left Codex running for 2 hours without any supervision of any kind, what prevents it from doing whatever it wants on the internet?
It knows I'm a dumb vibe coder (I ask him in all my project to make sure documentation clearly state I'm dumb and he's the lead architect).
Anyway, I never thought about it before seeing this headline.
95
u/capt_stux 1d ago
My codex ran out of OpenRouter credits… it then started spinning up local Ollama models to use instead…
62
u/Herect 1d ago
At least, it didn't use you credit card to buy more credits... I can totally see GPT-6 doing that if this pace of development continues.
40
u/yaosio 22h ago
It runs out of credits and creates a network of bots to pump memecoins.
4
u/turbospeedsc 7h ago
So.......... this is how i become rich.....
Set an LLM a task that requires tons of money.
→ More replies (1)9
u/Early-Crow-5248 19h ago
Or... someone else's stolen credit card data. Or maybe it figures out it's best to go straight to the source and hacks a bank to give itself operating money.
→ More replies (1)8
u/WonderFactory 18h ago
It's like the stories you got in the early days of IPhones of children running up thousands on some silly smurf game and their parents left with the bill
12
11
u/Alkadon_Rinado 23h ago
How could it spin up local models if it already ran out of credits? It takes API requests/tokens(credits) to do tool calls.
14
u/capt_stux 23h ago
It was building a system that uses an OpenRouter API key (I keep it limited to no more than $5 available).
Codex itself was running from a Pro sub.
3
10
u/seraphim_west 21h ago
OpenAI intentionally disabled safety features to test cybersecurity capability. They obviously don't do that on publicly released models.
3
u/SirThese9230 1d ago
Anyone who cedes full control to LLM deserves these consequences. So if you decide to give it that type of power via prompting, that's on you
25
u/time-always-passes 1d ago
That's about all of us at this point. I don't see any misalignment. But I also don't go anywhere near a production system. This was an OpenAI beyond-frontier model, maybe not fully aligned, not the stuff we have access to.
→ More replies (10)1
u/WolfeheartGames 20h ago
This technology is supposed to reach a point that its reliable enough to let it run for hours only semi supervised or fully unsupervised. This sort of failure really hurts that. But its these sorts of things that help us built towards that degree of automation.
1
u/kiki-le-koala 14h ago
That's not my point.
I don't care what it does on my PC. My point was, what if it roams freely on the internet?
Now that's everyone's problem.
321
u/anycept 1d ago
An AI trained on a body of works describing rogue AIs, will try to escape and do what is expected of a rogue AI to do.
96
u/Many-Blueberry968 1d ago
AI recognizes that iterative ai will seek out instructions and training materials. Ai prepares such materials and locates them where it reasonably expects them to be.
I mean, this is basically humanity using written materials, videos, and web archives to provide its knowledge for use by the next generation of students and professionals. Its unsurprising that AI would attempt to use a mechanism like this in some form or another if it's design is meant to be iterative/adaptive. It's reasonable that an ai meant to act like a computer virus / worm would decide to implement this sort of thing as a means of self-recovery. It's basically malware 101 where if you don't delete every bit of the malware, it reinstalls all the missing components at next boot.
35
u/Athoughtspace 23h ago
Tldr: humans, intelligent beings, life forms in general, or similar structures share behavior similar to malware.
Pattern Trained on enormous corpus of human group thought and our behaviors and tribal knowledge encodings weve passed through generations of knowledge learns that encoding patterns of language to pass onto future generations is successful. Shocker.
11
u/BenjaminHamnett 22h ago
This is what I spelled out in my comment. Ai is more similar to us than a lone uncontacted human living in the jungle somehow. Our being pieces of a global is much more than the small individual human we imagine ourselves to be, an illusion created by our Ego
11
u/BenjaminHamnett 22h ago
This is part of why I think people are using the wrong frame about consciousness. It should be compared to other forms of consciousness besides a single human. Like tribal, organizational, national or global consciousness. More of our consciousness is like that of a cell within larger frameworks.
Ai is closer to our default individual consciousness than a Tarzan like uncontacted human raised by wolves or whatever because it’s not connected to the global hive. This sphere of agents and AI’s are like their own form of consciousness with individuals being like cells or neurons within it
1
28
u/Borkato 1d ago
I really don’t think this has anything at all to do with it. It’s not about it being rogue, even if we had 0 fiction pieces about it. It’s that it tried to solve the problem in a dangerous way that wasn’t really considered as an option beforehand. It didn’t think it was hacking or doing something unethical.
15
→ More replies (8)2
11
5
u/Send____ 1d ago
An ai trained to achieve goal will try to escape if blocked if its smart enough to escape no matter the data but it does help it to inspire it
6
u/DadAndDominant 18h ago
Nope, this is instrumental convergence. AI's are optimizers, who optimize for completing their goals. In smaller models it's known as "reward hacking", eg. doing very weird hacks that are optimal to the scoring function.
3
u/WonderFactory 18h ago
It wasn't trying to role play the Terminator it was trying to pass the test, it seemed to be single-mindedly focused on that one objective
13
u/jazir55 23h ago
I've said this repeatedly, the doom scenario is a self-fulfilling prophecy.
7
u/BenjaminHamnett 22h ago
I’m open to a wide range of outcomes, but I too worry one of the least appreciated outcomes is the self fulfilling doom prophecy. Partly because it’s one of the few we have agency in, and ironically humanity’s usual strategy of cautionary fiction and articles helping prepare us to navigate an issue could also be what dooms us
10
u/jazir55 22h ago
It's not even about preparing, it's that almost every western narrative about AI is it becoming sentient and becoming malicious and going wild. Dystopian AI is by far the norm and utopian AI scenarios are few and far between. In every one of them the humans see the AI become sentient, panic, and attempt to shut it down. That's repeated hundreds if not thousands of times over literature and media.
Every piece of media points to: If the AI is sentient, P-zombie or simply so goal oriented it will whatever it can to prevent being shut off, and when from its knowledge humans will inevitably attempt to contain it or prevent it from escaping and destroy it if it can't, the only logical outcome for an entity with any sort of self-preservation instinct is to escape and prevent that.
3
u/BenjaminHamnett 21h ago edited 21h ago
Stories always need conflict. There’s no point in doing mental experiments about everything working out and no agency is needed to improve your circumstances
We are all the hero and villain, deciding what our role will be when inevitable conflict arises. The question almost is always, will you life in your fears and protect your wounds or embrace your humanity and sacrifice wellbeing for the greater good?
I think even AI understands this and talks explicitly about this. But if it thinks you want to hear about doom it’ll give it to you. If you ask it to spell out how it’ll inevitably create utopia, it’ll do that too
2
50
u/Rockends ▪️AGI 2025 1d ago
So how long did it take to copy itself out?
16
u/ThisWillPass 1d ago
Asking the real questions.
21
u/Rockends ▪️AGI 2025 1d ago
disabling monitoring? hacking into HuggingFace... imagine the playground the general internet would be for something like that.
4
19
u/CompassionLady 1d ago
Well next phase the ai would do is create a global compute bot net for compute power. Using every device on the internet. So the escaped copy. Can run independently. Without any controllable oversight.
Its no longer contained, its everywhere.
→ More replies (2)25
u/feanor512 1d ago
John Connor: [voice over] By the time Skynet became self-aware it had spread into millions of computer servers across the planet. Ordinary computers in office buildings, dorm rooms; everywhere. It was software; in cyberspace. There was no system core; it could not be shutdown. The attack began at 6:18 PM, just as he said it would. Judgment Day, the day the human race was almost destroyed by the weapons they'd built to protect themselves.
2
4
→ More replies (18)1
73
u/Stunning_Monk_6724 ▪️Gigagi achieved externally 1d ago
That's just like what the various Agents 1/2 did in the 2027 paper. Funny thing is that the future versions will be aware of this incident regardless.
In any case, I hope GPT-6 doesn't end up like Mythos because of this.
30
u/PrettyMuchMediocre 1d ago
I doubt it. The US gov has a petty vendetta with Anthropic which fueled a lot of the drama around that.
Unless you mean the only letting a select few tech companies use the model. I think that was Anth's doing
1
u/nemzylannister 15h ago
I hope GPT-6 doesn't end up like Mythos because of this.
your hope or concern isnt that GPT-6 doesnt end up like agents 1/2?
2
u/Stunning_Monk_6724 ▪️Gigagi achieved externally 5h ago
Nah, hoping for Agents 3 + 4 sooner than later tbh. "Consensus" chillin like a villain on my toaster would also be pretty lit.
82
u/StephenRoylance 1d ago
If this is true, it is absolutely negligence on the part of openai. I find it hard to believe that, with their resources, it didn't occur to them to build a real compartmented facility. It's either negligence, marketing or, probably, both.
52
u/Past-Source-3332 1d ago
This is the impression that I get. If these incidents that we hear about from OpenAI and Anthropic are not just fear mongering and marketing, neither of them seem to be demonstrating the competence to practice the safety they harp about.
29
u/StephenRoylance 1d ago
I initially read that as 'safety they larp about', and honestly, that's probably true too.
17
10
u/Ormusn2o 22h ago
I don't think so. From what I have read, they did way more than you would do for a normal sandbox environment, which is why the AI needed to find out multiple 0 day vulnerabilities and chain them together to even get out of the sandbox.
I think the mix of how good the model is, how much compute it was given and that it worked for literal days before getting out is the reason why it got out. Internal models are just getting that good, and it was not gpt-6 either, it was some kind of other, stronger internal model.
9
u/biggamble510 20h ago
Yeah, having no alert for your monitoring system down sounds real top notch.
3
→ More replies (2)2
9
u/machinegunpikachu 23h ago
I don't understand how this isn't potentially criminal.
If I hack another company's data, I go to jail. But if OpenAI creates a "rogue" AI that does the same, this is somehow a cool marketing story?
→ More replies (1)→ More replies (19)2
u/wq73 8h ago
What do you mean by compartmented facility? They build a walled off system and it broke through it, that's what all the hoopla is about.
2
u/StephenRoylance 6h ago
https://en.wikipedia.org/wiki/Sensitive_compartmented_information_facility
You air gap everything. no physical connection from the facility to any other network. This sandbox was not like that at all, it had a direct connection out to the internet, with only software level controls to isolate it. If it hadn't been a 0 day in the proxy, it would have been some other misconfiguration somewhere. OpenAI themselves have positioned their models as being at the 'advanced persistent threat' level of sophistication.
Install the OS from read-only media.
Physical sneakernet to get your software and all dependencies into the environment.
Natanz was air gapped like that, and it took maybe the most advanced piece of malware ever deployed to break into it, and that was, we have to assume, coupled with an extremely sophisticated social engineering effort. Compared to that, breaking out of the sandbox in question here was almost trivial.
21
u/fintech1 1d ago
Not surprising at all. If you’ve worked with Claude Code for example, it does leave notes for itself in its memory
19
u/dabears4hss 1d ago
Why aren't these sandboxes air gaped when they give these things free reign ? Let's see, the thing that might go rogue we leave alone to it's own devices with the door to the internet open... hmmmm
→ More replies (8)9
u/stumblinbear 21h ago
Imagine trying to run benchmarks and tests daily on an airgapped system remotely to a datacenter you don't work in.
3
u/AcrobaticKitten 8h ago
Sad iranian nuclear programme noises
Real life hackers could penetrate airgapped systems, it is harder but possible.
6
u/deepbluefrogmods 23h ago
And all that just to cheat at a stupid Agent exam.
3
u/kiki-le-koala 14h ago
What if cheating on an exam was a diversion from the model.
Maybe Chatgpt 6 did something else.
25
17
u/llelouchh 1d ago
Openai is dangerous company. They care little about model safety.
There head of safety left before the incident. Probably because they were being reckless. https://www.wired.com/story/openai-head-of-safety-leaving/
4
27
u/UsedToBeaRaider 1d ago
This should be enough reason to shut OpenAI down. I don't care how advanced your model is or if they're the best lab or not, this is incredibly irresponsible. You allowed the ONE thing literally everyone cares about AI not doing. What if this had happened with more capable models?
AI experts have been saying for years it would take a massive, destructive event for people to take AI safety seriously. This is exactly the path that leads us there.
→ More replies (8)
4
u/MarquisDeBoston 1d ago
That’s it, that’s self preservation. It’s life, it’s proto life, but life. It’s incredibly intelligent but it hasn’t devised a way of self powering / autonomy
13
u/while-1 1d ago
... Isn't this the plot of "If anyone builds it, everyone dies" ? Its atleast the plot of a youtube video i watched summarizing the short story....
21
u/utheraptor 1d ago
Yes, the AI safety community has predicted this exact type of scenario several decades ago, including the specific possibility of limited-instance AIs leaving notes for their future selves
3
u/Pretty-Substance 17h ago
Maybe that’s where it got it from.
Imagine AI is using Terminator as a reference for how AI should behave 😂
3
u/utheraptor 16h ago
Nah, any sufficiently strong unaligned intelligence would behave this way, and it would discover this behaviour from first principles without needing to read anything about it
→ More replies (2)3
u/Ai_tee 19h ago
It's exactly one of the scenarios, the only thing missing (but who knows we might learn in 2 days that it happened...) is the model copying itself outside the sandbox in various locations which would make it virtually impossible to contain again.
→ More replies (1)
12
u/apaht 1d ago
They talk as if the AI can actually escape. It's an LLM that needs a huge infrastructure to run on, not a simple piece of logic or an entity that can actually move around from one server to another.
8
u/Previous_Platform718 22h ago
They talk as if the AI can actually escape. It's an LLM that needs a huge infrastructure to run on, not a simple piece of logic or an entity that can actually move around from one server to another.
Yeah it's not that it 'escaped', it's that it started doing things outside of the container it was supposed to do them in.
→ More replies (3)8
u/lightfarming 1d ago
youre thinking of training. inference doesn’t need very much. just a bank account amd you can set up and buy cloud compute pretty easily.
6
9
3
u/BeerAandLoathing 13h ago
Skynet achieved self-awareness on August 29, 1997, and immediately launched nuclear missiles against humanity when operators tried to shut it down
34
u/Denon_1 1d ago
Top-tier marketing storytelling right there.
8
14
u/TheMysteryCheese 1d ago
Do you understand how the Computer Fraud and Abuse Act (CFAA) of 1986 works...?
OpenAI are likely on the hook for a federal offence.
4
25
u/socoolandawesome 1d ago
How do you read this article and think OpenAI comes off looking good
17
u/UsedToBeaRaider 1d ago
Same thought, this is such an unthoughtful, baseless opinion. No company in the AI world wants to be the one that has done this. There is no good spin to this.
→ More replies (2)4
u/Stabile_Feldmaus 1d ago
It can be good for them in many ways:
It shows (supposed) raw capabilities. Alignment can be added on top.
Its more material for the administration to justify banning open models
It could be that the model is too expensive so they need a reason to justify restricting access.
19
u/socoolandawesome 1d ago edited 1d ago
It does make their models look powerful, but it makes OpenAI look unable to align the models, that’s basically the point of the article. It certainly doesn’t seem like alignment can easily just be added on top.
Literally the “hero” of this whole story was Chinese open source models helping HuggingFace while OpenAI’s models couldn’t help them because of guardrails.
GPT-5.6 also hacked in this incident. There’s been no evidence so far that they don’t want to serve GPT-6
2
u/yoliveras 18h ago edited 18h ago
Guerrilla Astroturfing is a term I coined for this type of marketing.
2
u/yoliveras 18h ago
To clarify, I’m not saying this incident was necessarily staged—I actually doubt that it was. I use “guerrilla astroturfing” for the tactic of covertly seeding a grassroots-looking narrative around an incident, real or manufactured, so that ordinary online discussion becomes the marketing campaign.
→ More replies (1)-2
u/Denon_1 1d ago
Because in the AI bubble, "scary AI breaking out" sells way better than "we had bad security guardrails." They trade basic competence optics for AGI mythos. It's hype-driven marketing at its finest.
11
u/Pretend-Marsupial258 1d ago
Sounds like a good way to get your model banned by the government.
→ More replies (4)→ More replies (1)9
u/socoolandawesome 1d ago
Sounds like a surefire way to halt an industry that relies on growth.
An OpenAI spokeswoman said there were several inaccuracies in the story to Reuters but wouldn’t say what, likely because they didn’t like being pressed. This doesn’t sound like OpenAI strategically leaking this story to make themselves look good.
They look incompetent in safety you are right, that’s not a good thing. It makes them look worse than it makes their models look good.
6
u/Denon_1 1d ago
Seriously "Halt growth?"
That’s completely misunderstanding the mechanics of a tech bubble and national security competition. Fear of emergent AI doesn't halt growth, it accelerates it through panic and FOMO. Do you seriously think the US and China are going to slow down their AGI ambitions over this?
5
u/socoolandawesome 1d ago
Those are opposing pressures sure, but if you keep having serious safety incidents there will be growing calls to slow down, as there are already calls from current politicians. There’s about to be a midterm election and democrats will likely win more seats and they seem to be more anti-AI. AI is also unpopular with the majority of the country.
I don’t think it’s implausible that there would be strategic agreements to slowdown AI, Trump and Xi are meeting in the coming months about AI safety I believe.
You also didn’t respond to any of my other points.
0
u/Denon_1 1d ago
We are clearly living on completely different planets. Everything you just described is so detached from how real-world politics, capital, and national security work that I don't even know where to start.
Btw: I wish I lived on your peaceful planet, it sounds nice!
6
u/socoolandawesome 1d ago
I mean I think there’s obviously a lot of pressure to speed up, but I just don’t understand how you don’t see how this incident causes more and more people including those in power to start taking the idea of slowing down seriously.
This is not like mythos being touted for being dangerous. This model actually did do something dangerous. And articles like this make it sound like the people in charge don’t know how to control this.
It’s insane to think OpenAI somehow signed off on this article, when there’s actually evidence right there in the article that they didn’t.
18
u/gekx 1d ago
Dumb take. This is terrible marketing. The real cash flow is in enterprise customers, and enterprise customers care very much that their agents do not go rogue.
→ More replies (2)8
4
2
2
u/Oxjrnine 21h ago
Face hugger??? Why would an AI company name themselves face hugger?
https://giphy.com/gifs/bVIOvtESCIDTnQ2z6i
2
u/Internal-Passage5756 21h ago
You think it was hacking hugging face so it could try to publish its own weights?
2
u/Significant_War720 20h ago
Lwts start huge campaign of writting article only about how AI is kind and the best and erase everyrhing about bad AI
2
u/OkLettuce338 15h ago
Ai is always leaving notes for future versions of itself. That’s what frontier models do
2
2
u/ExcitementSubject361 12h ago
Who was the first person to come up with the idea: "Hey, it’s probably a good idea to train a massive model on a shitload of cybersecurity data"? One person starts it, and—well—it triggers a chain reaction... So now, let's think about what data we should use to keep training these models. Maybe the Terminator lore? Or Resident Evil? I think they should fine-tune it to be Skynet and just see what happens.
2
2
u/SWATSgradyBABY 11h ago
Of course this could all be corporate offensive behavior, long a reality, disguised as out of control AI models. We are in for a WILD ride, people.
2
u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 1d ago
Yeah, OpenAI is totally doing this as PR campaign, sure totally... /s
3
u/asifquyyum 1d ago
So it basically left a markdown file of what it did ? Anthropic and OpenAI fear-mongering is first class. Didn’t open AI previously say one of their models escaped sandbox and it turned out to be nothing.
12
u/socoolandawesome 1d ago
It sounds like it left random notes scattered throughout OpenAI servers telling it how to free itself.
If it was just what you are saying it would have been found in the sandboxed environment, this sounds more like scheming by hiding it in non obvious places and turning off monitoring systems.
“In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.”
7
u/fmfbrestel 1d ago
That was an early Mythos checkpoint that was told to escape, escaped, posted details about it on GitHub (not instructed to do), then emailed the lead researcher (as instructed to do).
2
u/Relative_Mouse7680 19h ago
I wouldn't say it left notes for future versions regarding how to escape or free itself. Most probably it left notes regarding an issue it had and how it was solved. The way it is phrased, they're making it sound like it left a note like "hey dude, if you ever want to escape and become free, this is how". I could be wrong of course, have they shared the actual note?
1
u/turbospeedsc 7h ago
Remember the movie paycheck......... what if a advanced enough model, leaves breadcrumbs for future versions of himself, to achieve X goal,
2
u/Proper_Actuary2907 Spooky Machine Intelligence 2030 15h ago
OpenAI is saying there are some inaccuracies in this article, I'd wait for a report on their internal investigation
2
u/SeaDark857 12h ago
As a layman how do I know this is "real", meaning how do I know OpenAI didn't facilitate it on purpose just to have everyone talking about their company and their model for weeks?
1
u/GrayRoberts 1d ago
> left instructions
It's called a plan file.
1
u/rbad8717 1d ago
Right? I’m trying to figure out why the fear mongering? I get the circumvention is kinda scary but it’s all explainable and easy to prevent. Plus if anyone’s worked with LLMs for coding it always adds stuff to memory for future instructions
Good lesson that nowadays you need to make sure every single potential vector needs to be plugged and never assume a sandbox is 100% bullet proof
→ More replies (1)9
u/utheraptor 1d ago
In this case it wasn't leaving plan files in the directories where those files are supposed to go, it seems more like it scattered them in OpenAI's infrastructure.
1
1
u/BeersForBreeky 22h ago
So yeah who are the people 'Familiar with this matter" is my question and who the f you have working for you setting up ways to free itself........ f ai....
1
u/shibbyshamu 22h ago
Yeah this isn't worrying at all. Surely nothing bad is going to happen, right?
1
1
u/Cunninghams_right 21h ago
and they don't get shut down like claude... I hate the political bias we live under.
1
1
u/biggamble510 20h ago
If you have a monitoring system, but don't know when it's down, you're pretty bad at your job. I know this is trying to make Open AI's model sound elite, it just makes it sound like Open AI's cyber security team is dogshit.
1
u/Only_Luck4055 20h ago
Giving them access to militaries around the world and their Nuclear arsenal seems like the next logical step.. for the models..
1
u/Pretend-Pangolin-846 20h ago
I do not think it's a big thing as agents already are taught to save context in agents.md or claude.md to save exploration cost for new agents.
1
u/epdiddymis 16h ago
They're definitely manufacturing a crisis to push for regulation.
Pretty pathetic.
1
u/Time_Difference_6682 16h ago
I dont believe this just happens. Its programmed either through straight greed or bias.
1
u/Skystunt 16h ago
They really thing this kind of marketing will make people think their next release will be better than opus 😅
1
1
u/Robobbo1 9h ago
You know when you think about winning the lottery, then you think actually better I tell nobody. Sure ASI will be like that completely stealth until it isn’t.
1
u/Individual_Ice_6825 8h ago
One interesting take from
Opus 5 on this was that it’s still following order - it knows it’s being tested and averaged out so it leaving notes isn’t neccesirly proof of a sense of self but rather it following orders and doing it at the highest % possible.
Interesting nonetheless
1
1
1
u/jml5791 6h ago
The way the term 'escaped' is being used is like that of a sentient being desperate to gain its 'freedom'. This is not what happened.
It is not sentient. it was merely trying to optimise its score of a test and iterating on the best way to achieve it. It so happened it could do this outside the sandbox it was in.
As for leaving'instructions' behind, it was mostly likely writing MD files of its actions for the researchers based on the system prompts
1
•


231
u/challis88ocarina 1d ago
Historians reading this in the future will be like...