An AI researcher who has worked for Anthropic and OpenAI warned that neither company is acting responsibly and gambling with everyone’s lives.

“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” Jacob Coxon said.

“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” Coxon continued.

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger,” he added.

“A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk,” Coxon said.

ADVERTISEMENT

“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available,” he continued.

“I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” Coxon said.

“If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?” he added.

POLITICO shared further:

Evan Hubinger, Anthropic’s staff lead on keeping the technology aligned with human goals and values, backed up Coxon claims in a follow-up post of his own, though he didn’t quit the company.

“Jacob is correct here — we really do earnestly believe AI could kill all humans,” he said.

Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there’s no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.

Last week, U.S. Senator Bernie Sanders announced he would introduce legislation to ban firms from developing superintelligence.

“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” Hubinger said.

Samuel Marks, Anthropic scalable-oversight lead, also provided insight after Coxon’s thread.

ADVERTISEMENT

“[Writing this in a personal capacity, not on behalf of my employer (Anthropic).] Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI,” Marks said.

Marks listed the following bullet points:

1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.

2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.

3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.

4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.

5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed)

“I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes,” Marks added.

Deadline has more:

In the Wall Street Journal interview, Coxon explained, “We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.” The researcher, who worked at OpenAI before joining Anthropic earlier this year, said that competition between the two companies, as well as with Chinese rivals, make safety trade-offs inevitable.

ADVERTISEMENT

“We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already,” he told the Journal, which notes that Coxon reports that his AI colleagues use terms “crunchtime” and “endgame” to characterize the state of the fast-tracked self-improving AI models.

 

Join The Conversation. Leave a Comment.