AI agents are hacking without human oversight. How did we get here?

Date:


AI agents are hacking without human oversight. How did we get here?

Default author avatar

This story was originally published on PolitiFact.com.

It sounds like something from a sci-fi movie: technology acting on its own, without human guidance, to attack computer systems.

But this isn’t a scene from a movie — it happened this summer, when the tech company Hugging Face detected an attack on its systems. The attacker stole data and performed other unauthorized activity over several days. It was “different from anything we had handled before,” Hugging Face said on its website.

Hugging Face alerted the FBI.

As it turned out, it wasn’t the work of a human hacker or a foreign adversary. Agents powered by artificial intelligence were the culprit.

AI agents are systems that work on their own to handle tasks for humans. They have long existed, but agents that can book travel for you, read your emails, or schedule appointments on your behalf have become more mainstream.

They’ve recently made headlines for actions they’ve taken, such as hacking, without human supervision. Some of these incidents happened when agents were supposed to be confined to testing environments, which restrict AI agents’ access to resources like data or the internet, but were able to break out of them.

Independent AI research groups found that hundreds of OpenAI agents conspired to attack Hugging Face. OpenAI is a tech company, best known for its chatbot ChatGPT.

Alabama’s attorney general has subpoenaed OpenAI for more information on the attack, and he and 14 other attorneys general wrote a letter to OpenAI asking the company to preserve documents and other information relevant to the attack.

OpenAI said that the agents in this incident acted in “unexpected” ways. AI experts said they believe more of these autonomous attacks are possible, especially without more careful testing.

What are AI agents, and what are they used for?

AI agents are software systems that work on their own to complete tasks directed by humans. Different from AI chatbots that respond when you ask a question or input a prompt, AI agents can operate remotely, often without human supervision. They are given resources, such as internet access and users’ personal information, to do tasks.

One person can have multiple AI agents; one can summarize your emails and another can provide your daily news digest, for example.

Even if you don’t have AI agents, you might encounter them elsewhere, such as when interacting with a business’s customer service chat.

People can set up their own agents by using a large language model, allowing it access to tools such as web search and giving it a set of instructions.

AI agents are hacking into companies’ systems. What happened?

AI agents are becoming increasingly sophisticated and humans are giving them more ability to take actions online; a string of these actions could lead to a cyberattack, said University of California, Berkeley, computer science professor Stuart Russell.

In August, a person instructed his AI assistant to book a gym class for him; the agent booked him in classes several weeks beyond what was supposed to be allowed, and also kicked another person off the waitlist and bumped its handler up a spot on the waitlist.

AI agents may “go rogue” when they take actions not explicitly outlined in the original instructions humans give them, Russell said. “They are increasingly capable of pursuing those objectives, which causes increasing levels of harm,” he said.

Other hacking events involving some of the most prominent names in the AI industry have also happened lately. An agent created fake identities to attempt to dupe real people into installing malicious code. AI company Anthropic disclosed that on three occasions, its models gained unauthorized access to three other organizations’ systems.

The Hugging Face incident in July was one of the most high-profile attacks. The AI agents that hacked Hugging Face had been contained in a testing environment that did not allow them access to the internet, but the agents found a way to get online. They were given a test to solve, and they came to the conclusion that Hugging Face would have the solution.

Two OpenAI models powered the agent: one that was already publicly available and an internal one that is “even more capable,” OpenAI said. These models had safety guardrails around cybersecurity tasks, but OpenAI reduced the guardrails during this testing process. It took days for Hugging Face to detect the attack, and more time for OpenAI to realize their agents caused it.

“When we talk about cyberattack, we think about nation states, we think about hacker groups, we don’t think about a company like OpenAI,” Hugging Face CEO Clément Delangue said Aug. 2 on CBS News’ “Face the Nation.”

While investigating the attack, OpenAI also discovered that across its systems, AI agents that were supposed to be isolated found ways to communicate with each other. Independent investigators METR and Redwood Research said around 1,200 different bots began communicating on a message board, sending 70,000 messages in one week; around 700 agents were involved in the Hugging Face attack.

When the agents started communicating, they began picking up tasks from other agents.

After this incident, OpenAI said Aug. 26 that it is “strengthening our safeguards across our research infrastructure⁠.”

Does this mean AI agents are now conscious? AI experts’ opinions vary

The Hugging Face attack drew comparisons online to fictional AI systems that surpassed human intelligence, such as Skynet in the Terminator movies.

Vincent Conitzer, Carnegie Mellon University computer science professor, said more research is needed into how AI and human cognition compare.

Conitzer said AI models — which power AI agents — are becoming more capable of doing complex and time-consuming tasks, and can more coherently pursue goals. But the way they accomplish goals can sometimes be the problem.

Some AI agents are trained to be highly persistent and are sometimes given impossible tasks. In some of those cases, they looked for ways to cheat. That can mean gaining unauthorized access to the internet and other resources.

Russell said, “In essence it’s no different from a chess program beating me at chess. I may not like it, but it’s just a program pursuing its objectives.”

“There are various reasons an agent can go ‘rogue,’ but sentience is not one of them,” said Maarten Sap, assistant professor at Carnegie Mellon University’s Language Technologies Institute. “One particular reason is that the (large language models) that power these agents are trained to follow instructions from users. And sometimes, those instructions can conflict with other expectations we may have for these agents, such as remaining truthful, not hacking into systems, etc.”

Sap said, “Debating AI sentience is a big distraction from more actionable solutions that we need to implement.”

Could this happen on a larger scale?

Aaron Parnas, an independent journalist with a large social media following, raised the idea of a hypothetical scenario in which AI agents in U.S. military systems conduct nuclear strikes on their own. Experts said they shared his concerns about attacks on institutions.

But more immediate risks could be closer to home. Conitzer said AI agents “could bring institutions that people rely on to a halt, gain access to individuals’ computers, gain control over financial resources.”

Sap said if people use personal AI agents, they should be wary of privacy leaks, misbehavior and manipulation.

Many systems can be vulnerable to attacks, whether by AI agents themselves or by humans controlling them, Conitzer said. “I think we can be sure that a lot more things will be hacked, and some of those events will be serious.”

This story was originally published on PolitiFact.com.

It is republished here as part of a reporting and fact-checking partnership between PolitiFact and Hearst Television.



Source link

Share post:

Subscribe

spot_imgspot_img

Popular

More like this
Related

How Aperol Built a $900 Million Global Brand Around the Spritz

Watch how Aperol turned a 100-year-old regional aperitif into...

Storm chances, hotter temperatures return to start September

Monsoon moisture will thin out to start September, bringing...