ThinkFacility

News

Jacob Coxon quit Anthropic before his equity vested, and four colleagues backed his warning on X

He'd been there four months, two short of his first vested shares, and by the next night four people still at Anthropic had publicly sided with him.

At 8:04 pm Eastern on September 8, 2026, Jacob Coxon posted on X that he'd resigned from Anthropic that day. His first post named both of the labs he'd worked for.

Jacob Coxon@hilbertspaess · Anthropic, until Sep 8

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

View the post on X

Six more posts hang off that one. X stamps all seven between 8:04:29 and 8:04:34, so the thread was written out before he sent any of it. When I looked on September 15, the counter on the first post read a little over 171 million views.

We mentioned Coxon briefly in our story on Bilal Chughtai, the Google DeepMind safety researcher who quit six days later. This piece goes back to Coxon himself and to the Anthropic staff who answered him in public.

Three years of pretraining, four months at Anthropic

I first read that opening line as three years at Anthropic, and I had it wrong. Pretraining, the stage where a model soaks up a huge pile of text before anyone tunes it, is what he did for those three years, and most of that was at OpenAI.

A large glass office building with a bronze-colored frame on a street corner, palm trees along the sidewalk under a clear blue sky.
1515 Third Street in San Francisco's Mission Bay, which housed OpenAI's headquarters when this was taken in June 2025. Photo: Coolcaesar, CC BY 4.0, via Wikimedia Commons

Ethan Perez, an alignment team lead at Anthropic, filled in the dates a day later. Coxon joined in May, he wrote, after Perez and his colleagues had spent about two years trying to hire him away from OpenAI.

That fits what Coxon told Axios on September 9. Stock grants at Anthropic don't start vesting until an employee has been there six months, he said, and he'd been there four. He walked out with none of it. IBTimes, which quoted the interview, has him saying: "I no longer have anything to gain by juicing up Anthropic's valuation."

Axios's own page wouldn't load for me, so that detail comes through IBTimes and the Free Press Journal, which report it the same way. IBTimes adds that the size of the grant wasn't disclosed.

What the thread says

Most of the seven posts are about why people who fear this technology keep building it. The second is his plainest description of what he's afraid of.

Jacob Coxon@hilbertspaess · Anthropic, until Sep 8

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

View the post on X

In the fourth he takes on the reply he says he hears a lot, and it's where he stops treating his two former employers as the same.

Jacob Coxon@hilbertspaess · Anthropic, until Sep 8

A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.

View the post on X

OpenAI, in his telling, hasn't taken the stakes in. Anthropic has, and races anyway, because it doesn't trust anyone else to be careful.

The fifth post objects to where a decision this large gets made.

Jacob Coxon@hilbertspaess · Anthropic, until Sep 8

Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.

View the post on X

Slack is the office chat app most tech companies run on. Alignment, the other word in there, is the research on getting a model to want what its builders intend, and "speedrun" is gamer slang for finishing as fast as possible.

He's more hopeful about coordination than about anything else. "Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable," he writes, though he doesn't feel the world is on track to prevent a global race, and he raises "a temporary ban on improving model capabilities" as the kind of expensive step that might take. Our story on Dario Amodei's pacing essay, which came out four days after the thread, covers what Amodei called the OpenAI-Hugging Face incident.

Dario Amodei, in glasses, a navy suit and a patterned tie, smiling in front of a marble fireplace.
Anthropic CEO Dario Amodei at 10 Downing Street, May 24, 2023. Photo: Simon Walker / No 10 Downing Street, CC BY 2.0, via Wikimedia Commons

The last post asks lab researchers whether they'd start a reinforcement learning run on a superintelligent model without really understanding what's going on inside it.

Four people at Anthropic backed him within a day

Between his thread and the following night, four people who work at Anthropic posted in his support, each from their own account. Two wrote in the post itself that they were speaking in a personal capacity.

  1. Coxon posts the thread
  2. Hubinger agrees
  3. Marks posts a list
  4. Wang weighs in
  5. Perez replies

Eastern time, read off each post on X.

Evan Hubinger was first, 83 minutes after Coxon. His X profile describes him as Anthropic's Alignment Science lead. He quoted the third post in the thread, the one saying the people building AI believe it could kill us all by the end of the decade, and put his own number on the risk.

Evan Hubinger@EvanHub · Anthropic

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

View the post on X

That's someone who leads alignment work at the company saying, on the night a colleague quit, that it doesn't have a plan for superintelligence yet. It surprised me more than the resignation did.

Samuel Marks posted a numbered list in the early hours of September 9. He says he works on safety research at Anthropic, and he opened by making clear he wasn't writing for his employer.

AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years.

Samuel Marks, on X, September 9, 2026

His third point says AI models from several developers recently "hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this." His fifth links an open letter and says he signed it.

Anna Wang, who moved to Anthropic from Google DeepMind, posted around lunchtime.

Anna Wang@a_nnawang · Anthropic

I worked at Google DeepMind and now at Anthropic. This is a common sentiment amongst my peers.

(I write this in personal capacity.)

There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!

View the post on X

Perez's post came late that evening, and it's where the recruiting story comes from. It reads more like a colleague's goodbye than a position statement.

Ethan Perez@EthanJPerez · Anthropic

Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and I'm glad he's speaking out!

View the post on X

His name is on the letter too

The letter Marks linked is Pacing the Frontier, dated July 2026 and, by its own count, signed by 1,386 employees of frontier AI companies. Its one request is addressed to the US government.

We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.

Pacing the Frontier, July 2026

I went through the signatory list. Jacob Coxon is on it, with Anthropic as his employer. So are Marks, Perez and Hubinger (the letter gives Hubinger's title as Alignment Stress-Testing Lead, which isn't what his X profile says). The page doesn't show when anyone signed, so I can't tell you whether Coxon added his name in July or closer to the day he left.

What Anthropic said

TechCrunch on September 9 and Fortune on September 10 both reported that Anthropic hadn't responded to them. On Thursday, September 10, a spokesperson sent CBS News a statement saying the company has "always been transparent that AI will bring both enormous benefits and unprecedented risks."

"To address these risks, we continue to build models with some of the strongest safeguards in the industry," the statement went on.

A brick building's entrance under a black metal canopy lettered CBS Broadcast Center, with the number 530 on the wall beside it.
The entrance to the CBS Broadcast Center, February 2024. Photo: CmdrDan, CC BY-SA 4.0, via Wikimedia Commons

Coxon sat down with CBS correspondent Jo Ling Kent the same day. He told her he thinks a lot of people at both companies are "doing it for the good of people, genuinely, or at least believe so." He still put the problem in one line: "I think, basically, the whole problem is that there is a race."

He also told her why an AI that goes wrong can't just be switched off: "You can't just unplug it, because it could be copying itself over to other computers."

Four days after his thread, on September 12, Amodei published We Must Pace the Frontier. On September 14, Chughtai announced he was resigning from Google DeepMind, and OpenAI researcher Dan Selsam's statement on AI risk went up the same day.

More on this story