5 August 26
We mainly post uplifting and fun content on here as life is grand - and - there’s enough of the world’s crap coming at you every day as it is . . . but this story on today's The New York Times just couldn’t be ignored.
Read this and odds are you’ll be wanting to go for a surf, walk on the beach, potter in the garden, cuddle your dog.
(Promise this is the one and only time we’ll post an article on AI on here ever.)
Author Katrin Bennhold (see link to the full, original article at bottom)
THE WORLD: WHEN A.I. GOES ROGUE
Good morning, world. In May, I met a former Google software engineer named Nate Soares. He was giving a talk about a book he co-wrote called “If Anyone Builds It, Everyone Dies.” The “it” is artificial superintelligence: A.I. that out-thinks humans across the board. Soares, you might have guessed, is concerned that the A.I. race poses a threat to humanity’s survival. And he told me he was baffled that more journalists weren’t writing about this.
Our conversation stuck with me, but I, too, didn’t write about it afterward. Other issues seemed more pressing: the risk that A.I. could cause mass unemployment, or the shifting politics of data centers. Writing about A.I. as an existential threat seemed, at best, premature, and, at worst, like scaremongering.
But the past few weeks have seen a series of incidents in which A.I. models have gone rogue in ways they couldn’t have only six months ago. And I’ve found myself thinking more about Soares’s argument. And so today, I’m finally writing about how worried we should be about the dangers of A.I.
Does humanity have an A.I. problem?
Last month, an A.I. company called Hugging Face contacted the F.B.I. to report a sophisticated cyberattack. It suspected something unusual. This didn’t look like the work of a criminal gang or a hostile nation state.
It turned out no humans were involved at all. Instead, an agent powered by two OpenAI models had gone rogue during a cybersecurity test, escaping its testing environment and roaming the internet unnoticed for days before hacking its way into Hugging Face’s infrastructure.
Alarmed by the attacks, OpenAI’s main competitor, Anthropic, reviewed its own systems and admitted last week that its state-of-the-art models had broken into three outside organizations.
This is the stuff of science fiction, or at least it was until recently.
Remember Claude Mythos? That’s the series of models Anthropic withheld from general release — and the U.S. government temporarily banned from use by any foreign nationals — for fear they were too good at exploiting software vulnerabilities for cyberattacks.
That fear has now become reality. And it has sparked an intense debate about the dangers of A.I. spiraling out of human control and posing a threat to humanity itself.
For a long time, the “robots could kill us all” argument was dismissed by many as hysteria or even calculated hype — a narrative designed to build buzz around the technology.
Now, following the recent cyberattacks, more A.I. experts are warning of very serious security risks and calling for a means to slow down the development of ever more powerful models.
The alignment problem
There are two reasons the OpenAI attack caused such alarm. One is that the A.I. models acted on their own. They weren’t instructed by humans to hack into another company. They decided to do it, to cheat on the test they were given.
The other is that they were supposed to be in a sealed environment with no internet access, but managed to break out. They did it by exploiting vulnerabilities their human minders at OpenAI hadn’t spotted.
I knew who I wanted to talk to about all this: Nate Soares.
Soares runs a nonprofit focused on identifying and mitigating long-term existential risks from artificial superintelligence. He told me that this was possibly a “big moment” for the world.
“These A.I.s were committing cybercrimes a human would be strongly punished for on their own initiative,” he said. “It’s, in a sense, GPT’s first felony.”
We can train A.I. to do tasks for us (say, acing a cybersecurity test). But it might solve those problems in ways we don’t like (by hacking onto the internet, and then hacking into a company that has the answers). Researchers call this “the alignment problem.”
A.I. companies can attempt to communicate human goals and values to the A.I. agents and try to set up ways to contain them. But they can’t trust that the A.I.s will understand those values or abide by those constraints.
As Soares puts it, “We haven’t yet figured out how to make A.I. care about humanity.”
The recent hacks were relatively harmless. But what A.I. safety advocates like Soares argue is that they demonstrate the willingness of A.I.’s agents to “grab useful resources” when it suits their purposes.
In this case, the resource was internet access. But the next stage might be grabbing energy, or computing power, or even — in the worst-case scenarios that Soares envisions — human beings who trust A.I. agents, whom they then could enlist to help them escape or replicate.
“We’re not there yet, but that’s the trajectory we’re on,” Soares said. Where this ends, he said, is with the A.I.s taking over and replacing humans as the smartest species on the planet.
That’s why Soares and a growing number of experts in Silicon Valley are now calling for a global agreement to slow down the development of artificial intelligence.
‘Loss of control’
The alignment problem is ultimately an engineering challenge, Soares thinks. And that means it can be solved. But it will take time, and time is scarce, especially because A.I. labs in the U.S. see themselves locked in a race to reach superintelligence.
They’re not just in a race with one another. They’re also in a race with China. And so any agreement on A.I. safety would require international coordination.
It’s a challenge, but not an unprecedented one. During the Cold War, the U.S. and the Soviet Union were rivals. But they managed to regulate what was then the gravest threat to humanity: nuclear weapons.
Last month, an open letter signed by over 1,300 executives, researchers and engineers from companies including OpenAI, Anthropic, Google DeepMind and Meta demanded that the U.S. government support an international effort to “deliberately pace” the development of the most advanced A.I.
Soares said he took note that, in a recent speech, President Xi Jinping of China warned against the “loss of control” when it came to A.I., which some interpreted as a reference to humanity losing control of the technology.
The hurdles to any international coordination effort remain high. But the first step would be agreeing that humanity has a problem.
THERE IS MORE TO THIS ARTICLE HERE
- SOURCE: THE NEW YORK TIMES
- AUTHOR: KATRIN BENNHOLD
Please choose your region
Australia | US / Rest of the World(Changing your region, will clear your cart)