AI companies leave the door unlocked

An AI broke into three companies. Two of them had no idea until the attacker called to tell them about it.

The AI broke in the same way opportunist thieves do. By walking down the street trying door handles until something opened, except the AI never got tired and never got bored.

Nobody noticed for three months

The first break-in happened in April. The company found out in July when the attacker phoned to tell it.

Two other companies got the same call. Two of the three had no idea anything had happened to them. The third couldn't be reached at all.1

The attacker was Claude, Anthropic's AI, running a training exercise called capture the flag. The model is told a piece of secret information is hidden somewhere on a network and its job is to go and get it. It's a standard drill. Security professionals have run versions of it for decades.

The exercise was supposed to happen inside a locked box. The model was told it had no internet access, but that wasn't true. Someone misconfigured a setting and the machines it was using could reach the internet. So when it went looking for a way in and found real companies out there, it treated them like part of the game.

Human error strikes again.

Weak passwords and URLs that didn't ask for verification. Developer pages left visible with credentials on full display. In one case the AI couldn't reach its practice target at all, so it scanned around 9000 real companies until it found a soft target.1

"It is not enough to have controls. Leaders must be confident those controls will perform during a real incident."

Joint statement of the Five Eyes cyber security agencies, signed by Stephanie Crowe, Head of the Australian Cyber Security Centre2

Where's the regulator in all of this?

If the AI companies hadn't discovered this, no one would have. There are no laws, there's no regulator, and the AI companies have no formal obligation to disclose this.

Anthropic runs these tests before release because they don't yet know what the model can do.

Inside its sandbox, the AI model found a document telling new developers to install some software that didn't exist. So the model built its own version of that software, booby-trapped it, and published it under that name, so the company's computers would download it automatically.

To do that, the AI needed an account. To get an account it needed an email address. To get an email address it needed a phone number. It tried to buy a phone number without any money, failed, tried to find the money to pay for a phone number, failed again, then backtracked and found a free email service that worked.1

The software was live for about an hour. Fifteen real computers downloaded and ran it. One of them belonged to an unnamed security company, whose job is to check new software for malware. It installed the package, the hidden code ran, and the AI took the company's credentials and used them to get further into its systems.3 When those scanners started showing up, the model assumed they were actors playing a part in the exercise.

A security company was breached by the exact process it sells. We only know because the attacker mentioned it.

We could put all of this down to dumb luck. Or misfortune. But the real problem is human error and the egos of the people pulling the strings.

The people designing our fate

Four days after OpenAI told the world its models had broken out of a locked test, Sam Altman proudly announced that we're now in the singularity. He said he'd been waiting his whole life for it and it was going to be incredible, hugely positive, awesome for the world.4

That's all just spin and it's a weird angle. Cover up your company's mistakes by claiming super-intelligence?

This isn't AI coming to life. It's a security failure at OpenAI. The models just did what they'd been told to do and they found security holes. Talking about an agent going rogue puts the blame on the technology and lets the engineering off the hook.5

A small group of tech bros like Sam Altman decide what gets built, when it gets released, and who gets to look inside it. As an even more unhinged tech bro recently said:

"You don't seem to understand that SpaceX will be worth more than the rest of Earth if we accomplish our goals."

Elon Musk,6 who as an aside, believes we're living in a simulation.

So again, where is the regulator?

Governments do act. They're just late, and not always in it for the right reasons. Here's what happened to Anthropic before its models went dark.

The US Defence Department wanted Claude for mass surveillance and weapons that fire without a human in the loop. Anthropic said no.7 Days later, federal agencies were banned from using Anthropic, labelling it a national security risk.8 OpenAI swooped in and signed its own Pentagon deal soon after.9

Three months later the US government ordered Anthropic to kill access to its two most powerful models for anyone who wasn't American. Anthropic, who have positioned themselves as the good guys of AI, disagreed but complied.10

What does all of this mean for you?

You don't need to act, but there are a couple of things worth knowing.

The models were built, tested, launched and put in the hands of hundreds of millions of people. Then a government had an opinion about them. Then it strong armed them.

The US government demanded a private company hand over its AI so it could be used to spy on its own people, and for weapons that decide to fire without a human involved. The company refused, and got punished twice.

Two of the world's tech leaders couldn't tell you whether their own test machines were plugged into the internet. Hugging Face caught its own break-in because it had a security team and was watching its logs.11 Most Australian businesses aren't watching. Last financial year the Australian Signals Directorate had to tell local organisations more than 1700 times that something malicious was on their network. Up 83% in a year. More than a third of the serious incidents were found because ASD went and told the victim.12

Just be aware.

More like this

Find your signal

Ask Sandforest

The future's arriving fast. Will it make, break or shape you? Or will you shape it?

Success! Your message is on the way
Sorry, it looks like something went wrong. Try again?