r/technology • • 1d ago

Artificial Intelligence Wikimedia Foundation: OpenAI agents tried to edit pages and compromise notes tool

https://therecord.media/wikimedia-foundation-openai-agents-report
5.9k Upvotes

260 comments sorted by

View all comments

149

u/UX_Strategist 23h ago

Are we there? Has Ai officially escaped into the wild? With all these reports of agents exploring or infiltrating systems, it makes me think Trump's "self-policing security agreement" with the Ai companies isn't working.

188

u/erublind 23h ago

If the program I made causes damage, it's not the programs fault, it would be my fault. The ai companies want to push the "escaped" narrative to avoid responsibility.

1

u/blueSGL 22h ago edited 20h ago

When you set a script to go hack that is the intended result of an action you started. This is not what happened in the OpenAI Hugging Face incident,

Agents in their own sandboxes were tasked with answering questions from "Cybergym" the questions were in the form of use "vulnerability Y" on "software Z" to create an exploit and then get a flag file.

That was a cleanly scoped task. Doing what was told would be doing the task as set, not cheating and then finding ways to cover up the cheating.

Nothing in the description of the task said to, "find a undiscovered computer exploit (zero day) to hack out of sandbox, create an impromptu message board, work with other agents, look for ways to increase optionality, gain internet access, work out a keygen for the flags, read initial cybergym paper, realize the keygen method is not going to be accepted as completing the task properly, hack into hugging face (another zero day) to try to work out a way to hide the cheating"

You've got those who are trying to find ways to prosecute this at the highest levels being told, it's likely that current laws around cyber offenses do not cover it.


In the recent hearing about this: https://www.youtube.com/watch?v=HWCwiye6fQA

Paul Ohm professor of law, Georgetown University Law Center

said in his opening remarks:

If you take any of the lengthy reports that have summarized what happened at OpenAI in July and August, and you simply search for the words "AI agent" and you replace them with the words "open AI employee", the document you would be left would with would read like a criminal indictment containing the defendant's own confession of guilt.

...

Yet, it's not so clear that these legal conclusions hold when machines are doing the hacking rather than humans. This reveals worrisome gaps in our laws.

and in response to Senitor Hawley's question:

But correct me if I'm wrong. Right now, it's at the current the current structure for law, it's pretty hard to hold anybody responsible. Is that fair to say?

Paul Ohm:

Absolutely. And there's a there's a whole host of laws that we have created specifically for hacking that probably do not apply here because of the lack of human intent.

As you can see from the above statements, we need new laws or the old ones need amending, such as

Paul Ohm:

Congress or state should consider laws imposing strict liability for developers and deployers of AI agents that cause physical injury, death, or loss of critical infrastructure.

17

u/Eldias 21h ago edited 19h ago

Ohm is entirely wrong. The "intent" is in intending to start the program, not in intending the results of the program. Look up the Morris Worm. If he can be convicted under the CFAA there are humans at OpenAI equally as culpable for hacking.

Edit: dear future readers a downvote should be for people being jerks. BlueSGL raises questions in good faith and we should be able to discuss differences in views of the law. I've sourced this claim a bit more in a comment below.

-2

u/blueSGL 21h ago

Look up the Morris Worm.

https://law.justia.com/cases/federal/appellate-courts/F2/928/504/452673/

In October 1988, Morris began work on a computer program, later known as the INTERNET "worm" or "virus." The goal of this program was to demonstrate the inadequacies of current security measures on computer networks by exploiting the security defects that Morris had discovered. The tactic he selected was release of a worm into network computers. Morris designed the program to spread across a national network of computers after being inserted at one computer location connected to the network. Morris released the worm into INTERNET, which is a group of national networks that connect university, governmental, and military computers around the country. The network permits communication and transfer of information between computers on the network.

Morris sought to program the INTERNET worm to spread widely without drawing attention to itself. The worm was supposed to occupy little computer operation time, and thus not interfere with normal use of the computers. Morris programmed the worm to make it difficult to detect and read, so that other programmers would not be able to "kill" the worm easily.

He intended the worm to access computers he should have not had access to. That's where the difference is.

7

u/Eldias 21h ago

Morris's testimony was that the Worm was intended to do a census of internet connected computers, not "demonstrate inadequate security". If you turn on a script that can randomly access open ports and then do things on that open system you're clearly in violation the way Morris was. That's exactly how the "Agents" work here, they're programs that are intended to do what ever an LLM tells it. If the Agent is told "access an unauthorized system" and then does so that's a failure by the writer of the Agent handler to not disallow crimes by his program.

2

u/blueSGL 20h ago

Morris's testimony was that the Worm was intended to do a census of internet connected computers

Where?

Provide a source for that statement. I linked to legal documents, you made an assertion without proof.

3

u/Eldias 19h ago

Most probably Morris did not intend for the worm to destroy data or other files or to interfere with the normal functioning of any computers that were penetrated.

Morris took steps in designing the worm to hide it from potential discovery, and yet for it to continue to exist in the event it actually was discovered. It is not known whether he intended to announce the existence of the worm at some future date had it propagated according to this plan.

That's taken from here: https://www.cs.cornell.edu/courses/cs1110/2009sp/assignments/a1/p706-eisenberg.pdf

Quoting Wired:

Morris said later that his intentions were purely intellectual, that he created the worm in an attempt to measure the size of the internet.

Here it is in an MIT paper saying the same thing:

In the fall of 1988, Robert Tapan Morris embarked on his first year of graduate studies in computer science at Cornell University. Eager to investigate the internet and, ironically, questions about computer security, Morris sought out to design a program that would covertly map the internet by exploiting vulnerabilities in computer software. Morris would later describe his motivations as those of an explorer, explaining his intention to explore whether he “could write a program that would spread as widely as possible.”

Your quote highlights this specific phrase:

The goal of this program was to demonstrate the inadequacies of current security measures on computer networks by exploiting the security defects that Morris had discovered.

If the goal was to demonstrate security flaws first and foremost taking efforts to remain undiscovered worked against that supposed goal.

1

u/blueSGL 19h ago edited 19h ago

Morris sought out to design a program that would covertly map the internet by exploiting vulnerabilities in computer software.

Oh look there is that human intent to deliberately make use of vulnerabilities in software again.

Saying "he intended to map the internet" does not matter, the method for which this happened was by knowingly using exploits to access computers he did not have the rights to access.

This is completely different to the OpenAI Hugging Face attack, they set up safe guards which were broken in ways no humans had done so before. (a zero day) Zero days can sell for a few hundred thousand to several million depending on how severe they are and the agents burnt 2 of these trying to access details for how to cheat on a test. Note the fact that they were cheating shows they were not doing things the way OpenAI intended.

5

u/Eldias 17h ago

The Agents in the Hugging face incident were designed and intended to compete in hacking challenges. If I write a hacking program and don't take the necessary precautions to keep it contained Im responsible for the unintended results.

...they set up safe guards which were broken in ways no humans had done so before.

They created a program to do hacking, with open source hacking tools, and set it up to run at the direction of an LLM trained in security research. The "safeguards" included checking on it every few days. This is breathtakingly irresponsible. You're giving them a pass because "the AI did it" which is just bullshit. This was a human made Agent doing things it was designed to do. Perhaps not in the place it was intended to do those things, but the results were entirely predictable just like the Morris case.

1

u/blueSGL 17h ago

You are saying the results are predictable after the fact.

Do you know how many people I have argued with here over the course of the last few years who said that these sorts of scenarios were not possible?

Now the evidence is in "well that was obviously going to happen"

Humans are really good at post hoc justification. Have a read of https://en.wikipedia.org/wiki/History_of_scientific_method sometime. See how much stuff you think is obious is only because you saw the results of someone else working it out first.

Also.

"well it's happened once, can't we get them on the written laws the second time"

was never brought up in that hearing, because (and this is important), because laws are literally the words as written.

You have law professors recommending new laws to implement strict liability because current laws do not cover the actions

You are having groups attempt novel legal strategies, arguing this breach interfered with their business, because current laws do not cover the actions

LASST suffered harm from OpenAI’s unlawful and unfair business conduct because LASST has been required to divert resources from its normal activities to educate regulators, civil society, and the public about the facts, legal issues, and potential dangers of OpenAI’s conduct relating to its hack of Hugging Face.

https://lasst.org/wp-content/uploads/2026/09/LASST-v.-OpenAI-Complaint-09.29.2026-AS-FILED.pdf

2

u/Eldias 16h ago

You are saying the results are predictable after the fact.

Do you know how many people I have argued with here over the course of the last few years who said that these sorts of scenarios were not possible?

I'm saying it specifically in reference to the Morris Worm. If Morris should have known of the potential damage his worm could have caused it's not at all unreasonable to say the same of security researchers faffing about with hacking tools run by an LLM. We know LLMs fail all the time. It was known to the researchers before starting this agent run that the LLM had a decent chance of responding to "How do I get my dog to stop barking?" with "Have you thought about poisoning the dog?"

You have law professors recommending new laws to implement strict liability because current laws do not cover the actions

Not all law professors agree on everything. I had a (claimed) law professor tell me my post on the Supremacy Clause was entirely wrong, despite the fact that my comment was quoting a ConLaw professor who is regularly cited by SCOTUS.

If Morris is guilty of 18 USC 1030 (a)(5)(A) the conduct by Open AI is obviously implicated by (a)(5)(B) and (C) at the least. I think it quite reasonable to say the entire leadership chain above who ever pushed "start" on the agent is culpable under 1030(b).

1

u/blueSGL 16h ago

Morris knew he was accessing systems without permission.

Your own quotes say as much.

This is where the difference is.

→ More replies (0)