When the person whose job is safety declares that safety cannot be achieved and walks away, is that betrayal or honesty? On September 9, 2026, a twenty-seven-year-old researcher in Anthropic's alignment division resigned. The WSJ, CNBC, and Forbes all covered it for the same reason: what he said on the way out. 'The probability that AI kills all of us within the next ten years is above ten percent.' Those words came from inside a company that has placed safety at the center of its identity.
01The Day the Guard Left His Post
From its founding, Anthropic has sold caution while competitors sold speed. That is precisely why it matters when a safety researcher leaves that company. It is as though a firefighter declared 'This building will burn' and walked out of the station. The people left behind cannot tell whether the fire is real or the firefighter's judgment should be questioned.
On the same day, the Financial Times reported that Anthropic had withheld its latest model from the UK's AI safety testing body. A company that flies the safety flag is avoiding outside inspection. Place the two stories side by side and you can see daylight between the banner and the behavior.
02What Does Ten Percent Feel Like?
How should we receive the number 'above ten percent'? More than one chance in ten that humanity is wiped out. If you imagine a six-sided die, one face is extinction. That is not a probability anyone should ignore. Yet the number comes with no disclosed methodology. What assumptions were made, what model was used, what time horizon was considered for partial versus total catastrophe: none of that was shared.
Stating a probability is, by itself, a warning. But without showing how the probability was derived, the audience is forced into a binary: believe or disbelieve. There is no middle ground to occupy, no parameter to adjust, no scenario to debate. The number floats free of its derivation and takes on a life of its own in headlines.
The same thing happens in material review. When someone flags a claim as misleading, the flag means little unless it specifies how a reader might be misled and how often that misunderstanding is likely to occur. A bare assertion of danger, unaccompanied by evidence, asks for trust rather than judgment. And trust, once requested without basis, is hard to rebuild.
03Between Whistleblowing and Simply Quitting
Is this researcher's action whistleblowing? Strictly speaking, no. He did not expose wrongdoing. He expressed personal fear and said he was leaving the industry altogether. Whistleblowing has a destination: a court, a regulator, a newsroom. Here the destination was 'society at large,' and no concrete corrective plan was offered.
Whistleblowing
Reporting organizational wrongdoing or danger to an external authority. Legal protections exist. A specific corrective action is sought.
Fear-driven departure
Leaving one's post because of perceived danger. No specific authority is addressed and no corrective responsibility is assumed, though the act of leaving can itself become a message.
Prophetic warning
Publicly stating a future danger. No obligation to present evidence. Difficult to verify. If wrong, no consequences; if right, acclaim. An asymmetric position.
In Eichmann in Jerusalem, Hannah Arendt wrote about the danger of people inside organizations ceasing to think. The reverse also occurs: someone who keeps thinking may eventually leave. The question is not whether the person who left was right, but how those who remain will think from now on.
04The Double Bind of Safety Claims
Building a corporate identity around safety creates a structural trap. Say it is safe, and you are told to prove it. Say it is dangerous, and you are asked why you keep building. Gregory Bateson called a similar structure a double bind: every move invites criticism.
| Position | Claim | Criticism |
|---|---|---|
| Claim safety | Our models are under control | Show evidence. Submit to external review |
| Admit danger | The risks are high; we proceed with caution | Then stop building. Stop selling |
| Stay silent | No comment | Lack of transparency. Something is being hidden |
On the same day, The American Prospect reported that Anthropic had been building a predictive surveillance system aimed at monitoring activists. A company flying the safety banner is developing surveillance technology. The wider the gap between banner and behavior, the faster trust in the banner erodes.
05A Familiar Tension in Material Review
Our own work carries a similar tension. Stamping a promotional piece as 'approved' is close to a safety declaration. But whether the person who stamped it truly checked everything is invisible from outside. Every box on the checklist may be ticked, yet there are places the eye did not reach. The stamp says 'safe,' but the stamp cannot speak to what it did not see.
When a reviewer blocks a piece with 'This cannot go out,' the refusal looks like fear if the reason is not articulated. Conversely, if the piece goes out and a problem surfaces, the question becomes 'Why did you let it through?' Those whose job is safety live permanently in this crossfire. There is no position that escapes criticism entirely.
That is exactly why the ability to put reasons into words matters. Not 'I am afraid,' but 'This passage is likely to be read as X because of Y.' Separating fear from evidence is the core of the reviewer's craft. The reviewer who can articulate the mechanism of misunderstanding earns the right to block. The reviewer who cannot is left holding a feeling that persuades no one.
06What to Ask the Person Who Left
If there is something worth asking this researcher, it is not 'Is ten percent really right?' It is: 'While you were there, what did you try to change, and what could you not change?' A probability is only the entrance to a discussion. The exit is concrete action.
In The Plague, Albert Camus depicted Dr. Rieux, who stayed in the stricken city and kept treating patients. Rieux was no hero. He was simply a person who chose to remain at his post. There are times when leaving is the right call. But before leaving, one is asked what was attempted.
What leaving conveys
The fact that speaking up inside the organization did not produce change. The departure itself can become the strongest possible message.
What staying protects
The ability to keep watching problems visible only from the inside and to accumulate small corrections. However, staying risks being seen as complicity.
07Turning Fear into Work
Those whose job is safety work alongside fear every day. What if I missed something? What if I approved something that should have been stopped? That fear does not disappear. Nor does it need to. Fear is a signal that tells you where to direct your attention.
The fork in the road is whether fear paralyzes you or whether you convert it into evidence and take the next step. Throwing a number like ten percent and leaving, and staying to chip that number down by a tenth of a percent at a time, are both forms of honesty. But the chipping is hard to see in the record, while the thrown number makes a headline. The quiet work of incremental safety improvement never trends on social media.
The single promotional piece sitting on our desk shares that same structure. No one can eliminate every risk. But risks can be shaved, one by one, with stated reasons. A sentence reworded so it will not mislead. A chart rescaled so the difference is not exaggerated. A footnote added so the limitation is visible. None of these actions will appear in the news. All of them reduce the probability, by some fraction, that someone will be harmed.
On the day the safety guard ran, the job of those who remain is to go back to their posts. Clutching their fear, refusing to let go.
Talking about safety and building safety are different things. Words make headlines; the work rarely does. The 'ten percent' thrown by the Anthropic researcher is valid as a question. But answering it is not the job of the person who left. It belongs to those who stayed. Converting fear into evidence, evidence into action. That quiet, endless repetition is what actually builds safety.
Key Points ── 3 to take away
- When fear emerges from inside a safety-first organization, it is proof that thinking is still happening there. The real question is whether the fear changed the organization's actions.
- A probability is only the entrance to a discussion. Without disclosed methodology and a concrete corrective plan, a warning remains an opinion.
- Fear need not be eliminated. Converting fear into evidence and shaving risk one piece at a time is the real work of anyone who carries the safety title.
Sources & references
- Wall Street Journal, "Anthropic Researcher Quits Over 'Uncontrollable' AI Concerns," 2026-09-09.
- CNBC, "Anthropic researcher says 'probability that AI kills all of us' is 'above 10%' after colleague quits," 2026-09-09.
- Forbes, "Anthropic's Head Of Alignment Warns AI Has 'Above 10%' Chance Of 'Killing All Of Us' Within Decade," 2026-09-09.
- Financial Times, "Anthropic withheld latest AI model from UK testing body," 2026-09-09.
- The American Prospect, "Anthropic and predictive surveillance systems," 2026-09-09.
- Hannah Arendt, Eichmann in Jerusalem: A Report on the Banality of Evil, Viking Press, 1963.
- Gregory Bateson, Steps to an Ecology of Mind, University of Chicago Press, 1972.
- Albert Camus, La Peste (The Plague), Gallimard, 1947.