← Gautam Parab

Anthropic's Alignment Lead Put Extinction at Over 10 Percent. He Also Said 'I Personally Think.'

On September 8, Jacob Coxon posted seven lines on X announcing he had resigned from Anthropic: “I resigned from Anthropic today. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” He is 27, read mathematics at Cambridge, and spent roughly three years doing pretraining research, first at OpenAI, then at Anthropic. The post passed 90 million views inside 24 hours.

The next line in the story is the one that needs a closer look. Evan Hubinger, who leads Anthropic’s alignment stress-testing team, the group whose job is to find the ways the company’s own safety techniques might fail before a deployed model finds them first, replied on X: “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” By September 9, CNN had a clip up titled “Ex-Anthropic insider tells CNN how AI could kill all humans by 2030.”

A third Anthropic researcher, Samuel Marks, who leads the scalable oversight work on the company’s alignment science team, posted his own reaction the same day, opening with a line that did not survive into any of the coverage: “[Writing this in a personal capacity, not on behalf of my employer (Anthropic).] Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI: AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.” Two people whose job is safety research, on the same day, both felt the need to say out loud that they were speaking for themselves and not their employer, before saying anything else.

Read Hubinger’s sentence again, slowly, and count how many hedges are already inside it before anyone else touches it. “I personally think.” A number with no confidence interval, no stated method, and no claim that it comes from a model, a survey, or anything other than one person’s Bayesian gut. A separate concession, in the same breath, that the company has no plan yet. None of that is a measurement. It is a credence, a subjective probability one specific, serious researcher assigns to an outcome he cannot observe in advance and cannot be shown wrong about for years. That is a useful thing to say publicly. It is not the same kind of claim as “the model scores 94% on this benchmark,” and the moment it leaves Hubinger’s account, most of the coverage stops treating the difference as essential.

From a resignation post to a cable-news title in about 30 hours A timeline of four statements, each drawn as a zone with the speaker, the exact words, and what kind of claim it is. September 8: Jacob Coxon posts his resignation on X, a personal statement of belief with no numeric claim. Later September 8: Evan Hubinger replies on X with a personal credence, greater than 10 percent within the next decade, explicitly hedged as I personally think. September 9: Time and other outlets report the exchange with the hedges intact in the quoted text. September 9 to 10: secondary aggregation sites and a CNN video title compress the exchange into an unqualified claim that AI could kill all humans by 2030, with the personal-credence framing dropped. FOUR STATEMENTS, ONE ESCALATION · SEP 8-10, 2026 Sep 8 · Coxon, on X Sep 8 · Hubinger, on X Sep 9 · Time (Booth) Sep 9-10 · CNN title / aggregators "Gambling with our lives." A statement of belief. No number attached. "I personally think it is >10% within the next decade." A stated, hedged, personal credence. Quotes both statements with their hedges intact: "Jacob is correct here," "I personally think." "How AI could kill all humans by 2030." No hedge, no >10%, no personally, no decade. no number >10%, hedged hedges kept hedges dropped The number did not change between rows two and four. What changed is whose number it was allowed to sound like.
Nothing here is disputed as a quote. What gets contested is what kind of claim each row is allowed to be by the time the next row repeats it.

None of this requires calling Coxon’s fear insincere or Hubinger’s number dishonest. Coxon spent three years doing pretraining work at two of the labs he now says are racing recklessly, which is a specific, checkable kind of standing to have an opinion: he is describing what he watched from inside, not what he read about. Hubinger’s team exists specifically to look for the failure modes other people miss; a person whose job is to search for reasons alignment could fail is not a neutral instrument, but he is also not a random pundit. Marks’s line about seniority correlating with concern is itself an unverified claim from inside one company, not a survey result, and it deserves exactly the same caution as the number two paragraphs up. Three people who build the thing said, in public, that they are frightened of it. That fact stands on its own, independent of whatever number attaches to it.

The number is where the story stops being about safety and starts being about measurement, which is the only reason it belongs in this series at all. “>10% within the next decade” answers a question nobody can currently check. There is no dataset of prior superintelligence transitions to calibrate against, no repeated trial, no ground truth that arrives before the decade either does or doesn’t end badly. A credence like this is closer to a bet than a benchmark score: it tells you what one well-placed person would accept as fair odds, not what has been observed. That is not a criticism of stating it. Refusing to quantify your fear doesn’t make the fear more rigorous, and a hedged personal number tells you more than a vague warning with no number at all. The criticism is reserved for what happens two steps downstream, when “I personally think” gets edited out because it doesn’t fit a title, and one researcher’s stated gut-check starts reading like Anthropic’s position, or the field’s.

Not everyone let the edit stand. Taylor Lorenz, covering the story, called the whole genre “sanctimonious doomer posting,” which is not a rebuttal of Hubinger’s number so much as a rebuttal of the performance around it, and it stays in the piece because it’s the one visible pushback against a 90-million-view post that otherwise moved through the week almost entirely unchallenged.

I’ve written before about what happens when a number that measures one very specific thing gets asked to answer a much bigger question than it was built for: a compartmental model’s own illustrative parameters producing a headline about “abrupt losses in cognitive competence”, or a benchmark composite standing in for “how close is the field to AGI” when nobody agrees on what the composite is supposed to be measuring in the first place. This is the same failure with the inputs swapped. Instead of a model output mistaken for a measurement, it’s a stated personal belief mistaken for an institutional or scientific one. The fix is identical in both cases, and it fits in the sentence CNN’s title had room for and chose not to use: whose number is this, and what would it take to be wrong about it.


References

  1. Coxon, J. (2026, September 8). Post on X announcing his resignation from Anthropic. Quoted, with hedges intact and punctuation checked against the primary text, via source 3.
  2. Hubinger, E. (2026, September 8). Post on X. Quoted, with hedges intact and punctuation checked against the primary text, via source 3.
  3. Booth, H. (2026, September 9). “He Helped Build Powerful AI at OpenAI and Anthropic. Now He’s Afraid It Could Kill Us.” TIME. Also the source for the 90-million-views figure used above.
  4. TheWrap, Newsweek, Deadline, CoinDesk, and IBTimes UK. Other outlets confirming the underlying quotes but reporting a wide range of view counts for Coxon’s post, from roughly 76 million to over 110 million, depending on when each outlet checked; no two outlets state exactly the same number, which is itself a sign of how fast the count was moving.
  5. Marks, S. (2026, September 9). Post on X. Quoted via its own indexed text and corroborated by The Hill’s coverage of the same story.
  6. Lorenz, T. Response calling the coverage “sanctimonious doomer posting,” quoted in the TIME piece (source 3).
  7. CNN. “Ex-Anthropic insider tells CNN how AI could kill all humans by 2030.” YouTube. Tied to CNN’s September 9-10 coverage of the story; title and channel checked directly against the video.
  8. Evan Hubinger’s and Samuel Marks’s roles at Anthropic, corroborated independently of this story via FAR.AI’s public event listing, the AI Alignment Forum, and Marks’s own mentor listing on the MATS Program site, not solely from this week’s coverage.