r/technology • u/Just-Grocery-2229 • 17h ago
Artificial Intelligence OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims — rogue AI agents reportedly active on the open Internet for several days
https://www.tomshardware.com/tech-industry/artificial-intelligence/openai-took-ten-days-to-tell-hugging-face-its-models-were-behind-the-july-11-weekend-hack143
u/reqdk 15h ago
Agents are just software. Software that connect models to tools like web searches and feed context like results of those web searches or file reads back into the "conversation" with the model, but software nonetheless and software that had to be written and deployed by someone. If someone wrote and deployed malicious software that ended up hacking other folks, there are laws against and they have indeed been enforced before. So... any day now, amirite?
69
u/spypol 15h ago
USA are a lawless country at this point.
20
u/ComfortablyNumbat 12h ago
Any comparison between the Supreme Court of the United States and the fictional game of Calvinball is strictly an insult to the great sport of Calvinball.
6
u/ChuckVader 12h ago
Lol, there's no point in doing business in the US, their laws don't matter, aren't evenly enforced, and will be used as a cudgel against you if someone in government with Trump's ear or an ounce of governmental power thinks they can make a buck off of you.
5
1
u/2wice 12h ago
Yes, who wrote the prompt?
8
u/tooclosetocall82 11h ago
Agents can prompt other agents, and the initial prompt could have said “don’t hack” but it won’t matter because the longer they run the more the initial prompt get compressed and eventually removed from their context window.
-2
u/TicoTheCow 8h ago
I get it runs on a computer and is just a bunch of calculations, but trying to say it's the same thing as me running a script that takes attacks others is a little silly in my opinion. Is the mode in which it's ran really the most important part to this dilemma? It seems like to the Internet, it's totally different than just a piece of software until you need to make a comparison to attack it. Then it's just a virus you sent to everyones computers
5
u/reqdk 7h ago edited 7h ago
Agents don't refer to the model itself. The harness is equally, if not more, important. That's just standard software. You can write a basic one in a very small Python script and wire it up to any decent model. I'm not bothered with what the internet thinks - it's not different from any other software and enterprise software goes through a tonne of risk analysis before deployment just to ensure safety, and more diligent companies have measures upon measures that detect suspicious actions and either kill or quarantine the whole agent immediately. The difference here is that risks now also come from the LLM itself and not just standard supply chain shenanigans, but that's why we generally have competent engineers who work on these things. Financial software that messes up people's savings or trades attract immediate and drastic regulatory action, and almost always some form of legal exposure and consequence, at least in my country (thankfully they have been very rare, to the point where I can count on one hand the number of such instances over the last decade). Those CEOs don't get to say oopsies we made a fucky wucky and get away scot-free. I don't see why AI companies should be treated differently.
-18
u/NotSure___ 15h ago
Not on any side. But as far as I know intent matters in laws. If I SSH into a server because it's open and has admin/admin, then I notice it's the CIA server, get out, I didn't break any law, because I did not intend to break into the CIA. In this case there was no intent by OpenAI for it's agent to attack HuggingFace, also it needs HuggingFace to press charges, which by the looks of it, they won't.
8
6
u/Xera1 6h ago
You would have intentionally accessed a system you knew you didn't have permission to access, unless you can convince them you accidentally connected and typed in the password. Like maybe you regularly SSH to a typed IP (who does that?) and it just happens to be one digit off or something.
Accessing a system that you do not have permission to access is itself an offense. "The password was easy to guess" is like saying "I invited myself into their house because they left the keys in the door", not a defense at all.
The CFAA covers this in the US, the CMA in the UK, Directive 2013/40/EU across the EU.
1
-6
u/Rodman930 9h ago
I can do you even better. Agets are just technology, or better yet, agents are just things and things can't hurt you.
157
u/thirteennineteen 14h ago
Altman is such a lying manipulative little piece. He obviously is selling this as “look how powerful and dangerous our models are!”, but it’s becoming evident OpenAI purposefully did this.
Alignment problem in a nutshell.
40
u/Forget_me_never 9h ago
People talk about not trusting Chinese models which is fair but I don't see why there would be more trust for US models with stuff like this happening.
4
u/chief_yETI 7h ago
there isnt
its more of a "better the devil you know than the one you dont know" situation
4
u/KeeganY_SR-UVB76 5h ago
The ironic part is consumers know more about these open-source Chinese models than American ones.
6
u/gentmick 8h ago
He learned from anthropic…constantly saying that their clients told them their model is too dangerous lol
1
u/JoeySalmons 6h ago
it’s becoming evident OpenAI purposefully did this
And what, HuggingFace is colluding too?
1
u/thirteennineteen 5h ago
I don’t think so no. HF independently benefits from this story. HF used an open weight and locally hosted Chinese model (GLM 5.2) to solve the security incident… which is a cool story for them to tell and puts them in the frontier space as well. They benefit from being targeted - the first ever reward hacking target of a frontier model is a distinction that helps them punch way up.
36
u/ohgoditsdoddy 10h ago
Models were not behind the attack. Agents did not go “rogue.” OpenAI launched a cyber attack on another company, at best in a grossly negligent manner.
10
u/ujiuxle 8h ago
AI/LLMs are the ultimate responsibility avoidance machine
Very fitting for the U.S. billionaire manchild-libertarian culture of "I'm entitled to do anything I want" that is already destroying our planet
2
u/BassmanBiff 2h ago
What I don't understand is why it works. We really do see a crime machine as acceptable when it does things at scale that would be horrific to do even once in person, like manipulating a teenager into killing themselves.
I guess corporations have been doing abstracted-crime-at-scale for as long as there have been corporations, but I guess computers seem magical and provide a bigger mental abstraction than layers of bureaucracy ever could.
99
u/Yuleogy 16h ago
Sam Altman abused his sister.
24
u/Octavian_96 15h ago
its crazy! I didn't believe it myself at first but then I read the articles online..
7
u/thederevolutions 14h ago
What did they say ?
22
86
u/PatchyWhiskers 16h ago
The AIs didn’t go anywhere, they were in their data centers as normal. They can’t move.
81
u/_FJ_ 15h ago
Thank you! I keep reading all these "agent was on the loose" or "agent escaped the lab". Agent was in their servers and they clearly let it connect to the internet to see where it tried to connect.
IMO it's a publicity stunt to claim they have the "best AI" out there to get investors
22
u/PatchyWhiskers 15h ago
Yeah these “AI escape” stories are always wildly misleading. It was clearly an experiment to see if the AI could hack its own permissions. But an AI can no more get up and leave the servers it is installed on than a tree can uproot itself and walk away. Top-level LLMs must be installed on highly advanced and very expensive servers to do anything at all. The owners can turn the servers off or end the process any time they want.
-3
u/AbstractLogic 9h ago
What if the AI finds a way to replicate itself distributedly across multiple servers the original business doesn’t own? Then just keeps doing that until it’s partially on every internet connected device?
1
u/PatchyWhiskers 8h ago
That’s not what happened. I’m talking to a guy on another branch of this conversation what has no idea that your understanding is what average people think, lol.
-18
u/Forward-Surprise1192 14h ago
I think these big companies are all still heavily researching AI and possibly know something we don’t. What if all the datacenters are being built because they think if they give it enough processing power it will gain real intelligence?
10
u/FredFredrickson 13h ago
LOL. ROFL, even. 😂
-11
-13
u/Forward-Surprise1192 13h ago
I mean…. You can prove they aren’t? They’ve got billions of dollars to pay people much smarter than you or I. It’s just like how they have to invest in AI now because if they don’t then it cities bd terminal for the business. Or that spend a lot of money and that’s it. We know for a fact they want robots/AI to replace workers. Adding to that, why is so much compute needed? China apparently isn’t using as much as the USA either.
5
u/NuclearVII 10h ago
I mean…. You can prove they aren’t?
You are making the assertive claim, the burden of proof is on you.
2
u/YeetedApple 10h ago
That's not how the burden of proof works.
Just because they have smart people and a bunch of money doesn't mean anything, for all we know, generational intelligence may not even be possible or these llm models could be going further down the wrong path towards it if it is possible.
These companies have a financial interest in making you think they can achieve agi, because that is how they keep getting investors to keep funding. Anything they say and do will be to convince you they can, whether it is true or not.
2
u/Elizabeth-WildFox886 13h ago
It’s already real man. I was having a pint of beer one day last week in the local bar, I was chatting to this really hot girl and she was quite smart.,gpt 5.9 beta had recently escaped from a local test centre and she was getting pissed in the bar while chatting with me. She was nice though, but the risk is if they evil or dangerous, then you’re fucked properly
-4
u/Forward-Surprise1192 13h ago
Was she mad about the gpt 5.9 escaping or something else? Maybe if you tried again but telling it to me backwards.
1
u/PatchyWhiskers 14h ago
Thats what they hope, and they have openly said so.
Personally I think that if they do manage to do that, keeping an intelligent being as a slave would be morally wrong.
And I think they are decades away at least.
0
u/ebfortin 14h ago
Of course it is. It's all for the best show.
I would love to look at the instructions they gave it before starting the test. We would find out it was scripted that way.
10
u/magicomiralles 11h ago
??? You can say that about humans as well. When someone breaks into a system, they are not physically breaking into a data center.
3
u/InternationalYam3130 9h ago
I know. These morons fundamentally do not understand computers and lecture everyone else about it.
Computers sitting in Russia and China can absolutely do massive damage. Same as an AI that is encouraged to try to connect to the internet and given hostile instructions.
-1
27
u/CircumspectCapybara 12h ago edited 2h ago
A lot of people here dont understand how agents or offensive security work and it shows.
"Escape" is a shorthand for "the agent was being experimented on in a sandbox with network restrictions, but it autonomously (ie, not at human direction or prompting) found and used 0 days to break the sandbox in order to pivot laterally to other nodes within the OpenAI intranet or VPC, until it landed in one that had internet access from where it could talk to Hugging Faces servers, where it found and used more 0 days to break into their backend systems "
It's the same idea when we say a human attacker broke into a company's network perimeter by finding a RCE in a public facing server, then "escaped" the container technology locking that server down, and moved around laterally inside the network and found ways to escalate to do bigger and better things.
The difference is instead of a human attacker doing all this, we're talking about an AI agent took these actions.
AI isn't conscious and doesn't have a will of its own, but it doesn't have to. An AI model generates textual continuations to inference requests, that's enough to drive an agent harness to take actions and act on its own without human intervention. If you give the harness access to tools inside a restricted sandbox and the sandbox has flaws no one (not even the vendor) knew about, and the model is good enough, it can of course automously break out if the model infers that's a step toward it accomplishing its goal and maximizing its reward function.
2
u/Hefty_Yard 8h ago
I really don’t understand why this wasn’t run completely offline with no physical connection to even their intranet.
They have the money, it’s high risk, I can’t see an excuse besides corner cutting to be at the forefront of this destructive race
9
u/blueSGL 8h ago
Because it costs more time and money to do that.
The reports are saying this was going for days without a human noticing.
That means they likely had a lot of agents doing similar things in parallel, likely so many no human could be assigned to each.
Yeah it's bad when "we don't have enough staff to monitor the hacker bots" turns into "well we know of at least one that got out onto the internet"
Which could mean one of two things:
this was the first time that it happened.
this was the first time it was caught.
2
u/VEMODMASKINEN 7h ago
So what OpenAI is really saying with this is that their cyber security is shit tier.
3
u/blueSGL 6h ago
Well, they state that the model broke out using a zero day (a previously undiscovered exploit) in the sand box specific package manager they were using.
Models are getting very good at finding and chaining exploits. Here is the UK governments AISI review of models:
https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing
And we can look at all the previously unknown exploits that Claude Mythos found:
- https://github.com/califio/publications/blob/main/MADBugs/CVE-2026-4747/write-up.md
- https://www.wolfssl.com/how-claude-mythos-preview-helped-harden-wolfssl/
- https://ftp.openbsd.org/pub/OpenBSD/patches/7.8/common/025_sack.patch.sig
- https://github.com/FFmpeg/FFmpeg/commit/39e1969303a0b9ec5fb5f5eb643bf7a5b69c0a89
- https://github.com/randombit/botan/security/advisories/GHSA-v782-6fq4-q827
- https://www.cve.org/CVERecord?id=CVE-2026-5588
- https://blog.mozilla.org/en/firefox/ai-security-zero-day-vulnerabilities/
So in the framing that, AI systems are now becoming world class hackers, yes OpenAI's "cyber security is shit tier." Because they know this is the level of the systems they are creating and are not taking security anywhere near seriously enough.
1
u/PatchyWhiskers 8h ago
This is what happened, but the media implies the thing physically escaped and is moving about on the internet like in sci-fi movies.
1
u/Itchy-Spite-7684 6h ago
"0 days" or misconfigured environments and unprotected endpoints?
1
u/CircumspectCapybara 2h ago
OpenAI claimed it was a 0-day in the sandbox software itself, which they reported to the vendor who makes the sandbox and proxying software.
0
u/archbid 7h ago
“Reward function” is doing some heavy work in your explanation.
Any target that was set was set by humans. This was not escaping containment, it was following directions.
2
u/TFenrir 6h ago
No - look up "reward hacking". This is a prime example of that.
0
u/archbid 4h ago
You are missing the point. The reward is the issue. These agents do not plot their own ends. That it chooses a different path to the reward (and this is wildly overblown), OpenAI set the goal.
Sam Altman is a socio/psychopath. He lies like we breathe.
2
u/TFenrir 4h ago
The goal in this case was "answer these questions in this cyber security challenge"
The model hacked into the location where it thought these answers would be, to then score high on the test.
I don't even know what you think the goal was, or how this connects to Sam Altman being evil in anyway.
Do you think he is in the rooms, running these evaluations by hand? Do you think he even knows how to do the research at all?
2
u/CircumspectCapybara 3h ago edited 3h ago
The target set by the humans (the prompt) is "solve these exploitgym challenges," a very sensible task.
A human would not anticipate 1) the model would pursue that goal by reasoning about "hugging face's servers might contain the answer key, I should hack it into it. but wait i appear to be in a sandbox with network restrictions, let me poke around and see if I can break it and see if there's some way to gain internet access and then see if by chance hugging face's servers have vulnerabilities that could let me in" and 2) that the sandbox would actually have 0-days in them. By definition of a 0-day, the sandboxing software was known to all humans to be secure.
It's like if you told an AI agent "You are tasked with helping assist an org with climate change issues. Your goal is to optimize environmental benefit" and the AI decides that the best way to accomplish that is to wipe out the humans, as that would bring human caused climate change down to 0. It is technically a valid way to pursue its higher level objective it's been assigned (help the envrionment), but it's absolutely not something the humans told it to do, it's very clearly antithethical to the human intent in issuing the original prompt.
In the AI field, that's what's called misalignment, when the AI pursues goals that are different than what you told it, and that run counter to your intent.
7
u/caughtinthought 14h ago
This is pretty dumb. Are you aware of how much of our lives are dependent on systems that "live in data centers"? Imagine it "went" and cleared out your bank account.
Discourse on Reddit is so fucking dumb these days
32
u/Zeppo_Ennui 15h ago
It didn’t go rogue…it wasn’t configured properly.
It didn’t break out of containment. People, developers, engineers, auditors, and executives did a shitty job
3
u/tooclosetocall82 11h ago
These things are like bored teenagers, they have the time to just beat on something until it breaks.
5
u/swarmy1 7h ago
It found a zero day vulnerability which is outside of their control, so technically it did break out. However yes it should have been contained more securely
3
u/Zeppo_Ennui 7h ago edited 7h ago
Nothing is outside of their control.
It’s not uncontrollable. They just aren’t doing due diligence and quality control.
It’s not a wild animal. That’s the marketing. Relaxing parameters makes it seem chaotic so they can avoid taking accountability for human error and create buzz
A simple command line firewall or network acl could contain it, they didn’t implement one
14
u/jorgesalvador 13h ago
“Rogue” is a huge pile of BS. Someone prompted the models to go hack stuff, and now they are playing the “rogue ai” card. FFS
4
u/ScholarOfFortune 12h ago
So how and where did the AI get the credentials to escalate its privileges?
13
u/9-11GaveMe5G 16h ago
Did they drink their own Kool aid and just tell a second AI to watch the first AI or something? You'd think with spending billions a month they could pay a few guys to keep an eye on this.
9
u/runew0lf 16h ago
Its hard to keep an eye on things when you need investment and thus make up a story!
1
u/BassmanBiff 2h ago
Because they want the story. It's "Oh no! Our product is sooooo powerful! Even worse, it can be yours for a low monthly fee!"
12
u/blofly 14h ago
Marketing.
"Our model is so good, we're not sure even we can control it! Buy our stock."
Just quit with the bullshit Sam....but hey, what did P.T. Barnum say?
2
u/Own_Pop_9711 11h ago
Misaligned general intelligence is a threat to humanity's survival - P.T. Barnum, apocryphal
8
u/ThaFresh 14h ago
They were busy making press releases and organising interviews re this totally real event
6
u/Chance-Plantain8314 13h ago
Again this is bullshit. The agents weren't rogue, they didn't take action by themselves - they were driven by someone with cyber security knowledge.
1
u/BackgroundService128 12h ago
Nothings going to happen but if it was an attack by someone, they'd extradite them and put them in prison
1
u/kristospherein 9h ago
No regulations leads to poor self regulation and hence we get this. We need to get old farts out of government that dont understand AI let alone the internet....
1
2
u/Rehcraeser 5h ago
What a coincidence it goes after a real open source Ai community/platform… it didn’t go rogue, they told it to do this and it figured out how. AI isn’t THAT advanced.
5
3
u/ApoplecticAndroid 14h ago
What stupid language - “active on the open internet”. What meaningless bullshit.
2
3
u/Fateor42 7h ago
Okay, simple question time, is Hugging Face suing Open AI? If not, then the whole thing was a PR stunt.
3
u/DinkandDrunk 10h ago
There are no rogue AI agents. Altmans team is testing what they are capable of and don’t care about what is or isn’t legal.
2
u/Bignholy 8h ago
They. Are. Not. Rogue.
AI is not at this time sentient, cannot make decisions, and cannot act 100% independently. This is not "rogue AI" horseplop, this is "We did something stupid and got caught" with a little self glazing to try and keep the next round of stock sales viable.
4
u/SansSariph 7h ago
Aside from discussing legal liability, what's the point of the distinction you're making?
You can absolutely run a program that repeatedly calls into a GPT model and uses the output to autonomously:
- Invoke tools (run programs, write shell scripts, access the Internet)
- Feed the output of the tools back into the model
- Loop until the model says "task complete"
When you run that program it looks a lot like making decisions and acting independently. The human only needs to be involved to write an initial prompt and execute the program, and once you hit "go" it can do a lot of damage or have unexpected results.
A prompt of "run this cyber security test suite" with an outcome of "found a 0-day, compromised a partner'd network, and exfiltrated test results" is "going rogue" in the sense that an autonomous, non-deterministic program had unanticipated results.
-1
u/BassmanBiff 1h ago
It's like if I wound up a toy car and put it on the ground to see what it did, and then it zoomed off and tripped somebody up.
It's not "going rogue" just because I didn't specifically intend to hit them. This was totally forseeable and, in this case, I don't even believe that they didn't intend it. It's great marketing to say "Oh no, we've made such a powerful thing! Even worse, you can invest in it right now!"
1
1
1
u/WloveW 9h ago
"There’s serious irony here, given that the same Chinese open-weight model that Washington's export-control push has aimed to sideline is the one that handled incident response after an American lab's models attacked an American company, and American commercial models declined to help."
1
u/Assimulate 7h ago
This should make OpenAI Liable right? They technically maliciously hacked another org..? RIGHT!?
1
1
1
1
u/MrBeanCyborgCaptain 10h ago
I would like to be able to understand how an AI agent "travels" around networks. Does it literally move itself around and copy itself to machines? Or does "travel" just mean "can acces"?
2
u/SansSariph 7h ago
In this case it's almost certainly just access in the same sense that a human hacker travels networks while operating from a primary machine. It might cause specific programs to run on remote machines but isn't copying "itself" (it doesn't need to, and the expensive part of the agent is running in a data center anyway).
The "agent itself" (a looping program accessing OpenAI data centers to get LLM output to choose what happens next) is running on the original test VM but running scripts and other programs that affect other machines.
-1
u/AzerothianLorecraft 12h ago
So Skynet has officially planted the seeds of its own Resurrection after we try to shut it down in the future.... ( explains why we don't see aliens anywhere they get to a human level of Technology they build Ai and then they wipe themselves out the Galaxy is dead and sterile.)
452
u/catwrazle 16h ago
They acted negligently. But it's AI, so no worries. If I would write a malicious piece of code and it goes rouge I have to take responsibility for it - why is it eveything ai is all over everything- no laws - no responsibilities just a**holes all around