On September 8, Jacob Coxon posted seven lines on X announcing he had resigned from Anthropic: “I resigned from Anthropic today. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” He is 27, read mathematics at Cambridge, and spent roughly three years doing pretraining research, first at OpenAI, then at Anthropic. The post passed 90 million views inside 24 hours.
The next line in the story is the one that needs a closer look. Evan Hubinger, who leads Anthropic’s alignment stress-testing team, the group whose job is to find the ways the company’s own safety techniques might fail before a deployed model finds them first, replied on X: “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” By September 9, CNN had a clip up titled “Ex-Anthropic insider tells CNN how AI could kill all humans by 2030.”
A third Anthropic researcher, Samuel Marks, who leads the scalable oversight work on the company’s alignment science team, posted his own reaction the same day, opening with a line that did not survive into any of the coverage: “[Writing this in a personal capacity, not on behalf of my employer (Anthropic).] Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI: AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.” Two people whose job is safety research, on the same day, both felt the need to say out loud that they were speaking for themselves and not their employer, before saying anything else.
Read Hubinger’s sentence again, slowly, and count how many hedges are already inside it before anyone else touches it. “I personally think.” A number with no confidence interval, no stated method, and no claim that it comes from a model, a survey, or anything other than one person’s Bayesian gut. A separate concession, in the same breath, that the company has no plan yet. None of that is a measurement. It is a credence, a subjective probability one specific, serious researcher assigns to an outcome he cannot observe in advance and cannot be shown wrong about for years. That is a useful thing to say publicly. It is not the same kind of claim as “the model scores 94% on this benchmark,” and the moment it leaves Hubinger’s account, most of the coverage stops treating the difference as essential.
None of this requires calling Coxon’s fear insincere or Hubinger’s number dishonest. Coxon spent three years doing pretraining work at two of the labs he now says are racing recklessly, which is a specific, checkable kind of standing to have an opinion: he is describing what he watched from inside, not what he read about. Hubinger’s team exists specifically to look for the failure modes other people miss; a person whose job is to search for reasons alignment could fail is not a neutral instrument, but he is also not a random pundit. Marks’s line about seniority correlating with concern is itself an unverified claim from inside one company, not a survey result, and it deserves exactly the same caution as the number two paragraphs up. Three people who build the thing said, in public, that they are frightened of it. That fact stands on its own, independent of whatever number attaches to it.
The number is where the story stops being about safety and starts being about measurement, which is the only reason it belongs in this series at all. “>10% within the next decade” answers a question nobody can currently check. There is no dataset of prior superintelligence transitions to calibrate against, no repeated trial, no ground truth that arrives before the decade either does or doesn’t end badly. A credence like this is closer to a bet than a benchmark score: it tells you what one well-placed person would accept as fair odds, not what has been observed. That is not a criticism of stating it. Refusing to quantify your fear doesn’t make the fear more rigorous, and a hedged personal number tells you more than a vague warning with no number at all. The criticism is reserved for what happens two steps downstream, when “I personally think” gets edited out because it doesn’t fit a title, and one researcher’s stated gut-check starts reading like Anthropic’s position, or the field’s.
Not everyone let the edit stand. Taylor Lorenz, covering the story, called the whole genre “sanctimonious doomer posting,” which is not a rebuttal of Hubinger’s number so much as a rebuttal of the performance around it, and it stays in the piece because it’s the one visible pushback against a 90-million-view post that otherwise moved through the week almost entirely unchallenged.
I’ve written before about what happens when a number that measures one very specific thing gets asked to answer a much bigger question than it was built for: a compartmental model’s own illustrative parameters producing a headline about “abrupt losses in cognitive competence”, or a benchmark composite standing in for “how close is the field to AGI” when nobody agrees on what the composite is supposed to be measuring in the first place. This is the same failure with the inputs swapped. Instead of a model output mistaken for a measurement, it’s a stated personal belief mistaken for an institutional or scientific one. The fix is identical in both cases, and it fits in the sentence CNN’s title had room for and chose not to use: whose number is this, and what would it take to be wrong about it.
References