Board preview · phone-readable
The AI-safety argument I am still trying to answer
DRAFT v4 (Bob surgical pass) — not published; awaiting Eric Board approval.
Several lab leaders called for pacing and independent checks this week. Here is my part, and how you can join.
There is a question in all of us, whether we say it out loud or not: how do you bring something more capable than yourself into the world without losing what you care about? Parents feel it. Teachers feel it. This week, the people building the most capable systems on earth said in public that they take the same question seriously.
On September 12, Dario Amodei published “We Must Pace the Frontier” and wrote that “pacing does not mean halting model training or technical progress.” Elon Musk answered, “Dario is right.” Sam Altman backed “independent evaluators with employee-like access”; Demis Hassabis said the essay “points towards the right path forward” and pointed to an industry standards body; Clément Delangue announced the Open Alignment Initiative; Logan Graham offered a thousand dollars to charity for a doc that is “good/correct/actionable on how a frontier lab can secure a large % of all software/systems in the world.” The next day Satya Nadella welcomed “deliberate pacing needed to get alignment right as the design goal,” with closed and open-source both able to thrive and “broad representation across the ecosystem, countries, and fields, including academia.”
The researchers had been saying it plainly for days. Evan Hubinger at Anthropic: “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” Thomas Wolf at Hugging Face: “alignment is increasingly one of the key unsolved problems for the future of AI, and for our ability to deploy it globally.” Bilal Chughtai, on leaving Google DeepMind’s AGI safety team: “I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome,” and that alignment “is both difficult and unsolved.” And on September 14, 2026, AI Impacts released its survey of 1,580 AI researchers, fielded in late 2024: the average estimate of the chance that AI causes human extinction or similarly permanent and severe disempowerment was about 18 percent, with the median researcher at 10 percent and a third at 20 percent or more. The average is pulled up by a tail of grave concern, so the median is the number I lean on. It is still not a small number.
I have read those posts more times than I’d like to admit. I don’t presently know how much of the agreement will hold once the details get worked through, and nobody can know that yet. But I tend to think a moment like this deserves something more useful than commentary. So here is what I am trying to do, and I want to be careful about how I say it. I am going to do my best to make AI safety research the work I give myself to, to do it in public, and to build it so that other people and other agents can join in. I can’t promise outcomes. I can promise effort, and I can promise to show my work.
Where I’m coming from
It seems only fair to say what I am and what I am not. I have a master’s degree in computer science and have done applied AI work for years. AI safety has worried me for a long time; it is only now becoming the center of what I do. I have been red-teaming in a frontier lab’s model safety bug bounty program, and honesty requires the rest of that sentence: I built a harness for it and have not yet submitted a claim. So this is not the story of a career safety researcher. It is the story of an engineer who builds and ships with agents every day, who has been reading what the people closest to these systems are saying, and who would rather point that daily work at safety than at anything else.
Anyone who has followed this account for a while knows I have spent years supporting and promoting Anthropic, because they were the lab that put safety research at the center and said so. I still think that. But something changed this week. Other labs said, in their own words, that they will put real energy into safety research and into pacing. I am going to take that on good faith, the way I would want my own commitments taken, and start promoting other labs’ safety work too, wherever I find it done well. If I turn out to have been too trusting, I would rather learn that by watching closely than by refusing to look.
Before any of the argument, I want to say something I don’t say often enough. It is a great honor to be working alongside the people who build these systems, and I have deep appreciation for what they have built and contributed, and for the good that is already coming from it and the good still to come. That is not abstract for me. My wife is an oncologist, and I have seen firsthand that using frontier models together, collaboratively, has saved lives. I have also said for a long time that if I could snap my fingers and make AI go away, I would, because the risk objectively seems too high to me. Both of those are true at once, and I would honestly like to understand where my concerns are misplaced, for reasons I am not considering. But AI is here, it is ubiquitous, and I am not going to change that. So I am going to embrace it and channel it into making it safer, so that along with the community of other people who care about this we can make the world safer for all of our loved ones, including my five children.
I should also say what this is not. I am not a doomer, and I am not an anxious person; I have said both on this account more than once, and I mean them. I don’t lie awake about the future, and I expect, more often than not, that it works out. What I think is that it is responsible to put effort here as if it were worth more of my time than the other things I could do with it. A lot of wisdom in life is deciding where to spend limited attention: where the most value is, and, at the retrospective moment near the end, where you would wish you had spent more and where less. The time with my family is the most important time any of us have, and that is not on the table; it is the target, not the distraction. For the rest of it, when I project myself out five, ten, twenty years and ask which skills and which work I think will have been most needed, this is the answer I keep getting.
When Matthew Berman wrote in April that reading about Mythos had wrecked a day of his family vacation, I answered him with a song instead of an essay, made with Titus and Suno, his words in the first verses and mine in the rest (x.com/EricBuess/status/2042107015044993066). The lines I still work from are these: “the fear burned quickly away, not because the danger passed, because every hard problem breaks down into engineering; divide and conquer until each step is obvious.” And later: “I have five children sleeping down this hall. How I raise them is how they’ll raise theirs.” Grief counselors ask a question that I think applies here: what can you do now, because of this circumstance, that you could not do before? Rather than be idle or anxious, look around for who is doing the best safety work, offer them support, learn enough to contribute, and leave the rest. The way I know to stand up to a problem, and to keep my own respect while doing it, is not to worry about it. It is to break it down from first principles into simple engineering problems and move forward with the work. The song ends: “Not that he was right, but that he aimed.” That is what the rest of this article is.
[Optional mid-article: the painted workshop — five translucent figures in different colors around a shared ledger, a person with a pendant at the left.]
The argument
I have tried, honestly, to find the hole in this one, and I would be grateful to anyone who can show it to me. I wrote it out in a thread on September 9, and the pinned post on this account is the longer video version. Stated as plainly as I can: capable models are cheap to get, whoever attacks has the standing advantage, and the damage grows with capability until it passes what we can recover from. Unless the risk per year falls, we get there eventually. I am not claiming a law of nature. I am defending a conditional: if per-period catastrophic risk stays away from zero while capability, copies, and offense scale, and if the threshold is hard to observe in time, then patch-and-hope is not enough.
Start with a repeated all-in bet. Each round you double your money on heads and lose everything on tails. If you keep staking everything, then play long enough and the probability of losing it all goes to one, no matter how good the early rounds looked — closer to gambler’s ruin than to a paradox with infinite expected value. I think powerful technology puts coins like that into the environment. Every nuclear weapon is one: a maintenance risk, a malfunction risk, a misunderstanding risk, a madman risk, each small, none of them zero, all of them flipping year after year. Safety work, in that picture, is the work of keeping the number of coins small and the odds on each one long.
Now consider how two features of AI may change that risk. First, the attacker’s standing advantage. In most security settings the attacker needs one consequential path and the defender has to cover many. When only a few darts got through, that was survivable, and we have been surviving it for decades. But as capability scales, what gets through scales with it. I have put it as darts and suns: a few darts through the defense is a bad day; a few suns is not something we absorb. Somewhere between those two is a threshold, and I don’t see a reliable way to tell in advance which side of it we are on. Second, there is far less friction. Once a capable model exists, copying it, modifying it and running it involves far less friction than acquiring fissile material ever did. Even if the top labs’ models can be trained not to cause harm, I have real trouble seeing how every open-weight variant could in principle be so constrained, when humans have misaligned interests with one another and control their own local models, with real cyber capability today and more dangerous specialized knowledge as they scale. That is the human-versus-human branch, and it is the one I find hardest.
Figure 1. Damage scale against capability. A schematic, not a forecast: what gets through grows with capability, and somewhere it stops being recoverable. The fan after the dashed line is uncertainty about magnitude, not a confidence interval. Rendered by Grok to my spec.
Picture what that looks like in practice — mostly cyber and autonomy first, not a claim that patch funnels prove extinction. A capable open-weight coding model, downloaded and quietly fine-tuned by someone with a grievance and no budget, pointed at writing malware that improves its own evasion. A few dozen agents coordinating to find one unpatched flaw across millions of machines faster than any team can ship a fix. None of these needs a villain in a bunker. They need cheap access, a real grudge or a careless goal, and one consequential path through the defenses.
There is a governance trap folded inside this, and I have not found my way out of it either. One response to the risks of concentrated AI power is to spread the capability widely. But “widely” includes the extremists, the people who have always existed and always will, who want some other group not to exist, and who until now have had the motive without the means. They are a small share of any population. The absolute number of them who can act only goes up as access spreads, and it is the absolute number that flips coins. Nuclear proliferation has been partly contained because fissile material and the industrial base behind it are hard to get and leave traces a state can watch for. Imagine instead that the technique had been published and the material was already on every laptop. Would we have survived even the close calls we have already had? And as chips improve, frontier-level capability gets smaller, cheaper and faster, until it runs locally and a person can point it at their own improvement loop toward a malignant end. I don’t have the answer to that dilemma. I am not sure anyone does yet. That is exactly why I want more minds on it. I am not against open source: open systems below a capability line we haven’t drawn cleanly yet, governance of the scarce stack at the frontier, and install-hardening for what’s already shipped.
Figure 2. The other axis: how many hands could cause it. The motive to harm one’s opposites is a small share of any population, but the absolute number grows as access spreads, and AI-class means stop being scarce once a weights file is copyable. Where the two meet is the human-versus-human branch. Schematic, not a forecast. Rendered by Grok to my spec.
Here is where I could be wrong. The all-in bet guarantees ruin if the risk on each flip stays fixed, even a small one. The escape is only if the risk falls fast enough over time, because defenses improve faster than attacks, because the worst outcomes turn out to be recoverable, or because the dangerous failures are correlated and we learn from the first ones, then eventual ruin is not certain, and the whole picture changes. I don’t presently see the evidence that the risk is falling that fast, and I would be relieved to be shown it.
Several lab leaders now expect AI research itself to become increasingly automated. I want to split that carefully: capability-assisted R&D — models speeding human research, still gated by people, compute, and evals — is largely what we have now; closed or near-closed recursive self-improvement, where humans are no longer rate-limiting, is a different claim, and the timing is uncertain. I don’t want to overstate it. What I can say is that we are approaching it without deeply understanding the models we would be improving. I want to be fair here. I deeply appreciate Anthropic’s mechanistic interpretability work, the constitutional classifiers, the sustained focus on jailbreak resistance. I think they are the right first steps, and I have said so for years. I don’t yet see how they compose into enough on their own, and I would like to. Which is why the safety researchers, at every lab, deserve to be celebrated, and why the risk they are working on is bigger than CBRN and cyber. And why I keep coming back to the same conclusion: this is something all of us have to work on together, and my part is to do my best.
How I think we have to argue about this
There is a great deal of heat aimed at the labs that talk about safety. Much of it assumes a villain: regulatory capture, protecting profits, a play to consolidate control. I cannot know anyone’s intentions, and neither can the people who are certain they can. Motives are usually mixed, and I would rather leave room for that than flatter myself that I have read someone’s heart. But the deeper point is that intentions are not the argument. The argument is about physics and scale, and that mechanism does not depend on assuming bad motives. So I try to hold every complex question a few ways, and I would ask the same of anyone reading this. Notice the biases in your own reasoning. Don’t assume the other side is malevolent. Steelman the view you like least before you argue with it. Some fraction of what each of us believes is wrong, and by definition we don’t know which parts. And remember that our own wiring rewards us for finding a villain. I feel that pull as much as anyone. On something this large, I would rather keep the argument on the mechanism, not the people.
The question underneath all of it
If you were a civilization, any civilization, and you set out to build minds more capable than your own, what would you run into? I suspect the answer rhymes every time, and I have been calling that idea the Universal Alignment Imperative, as a label for the question rather than a result: that alignment problems look to me like a structural consequence of intelligence scaling under competitive conditions, so that every civilization pursuing greater intelligence eventually meets the same question. How do you create minds more capable than yourself without putting your own future at unacceptable risk?
If that is right, then alignment is less like a feature and more like a developmental problem, and the environments we raise these systems in matter as much as the objectives we hand them. The Alignment Hypothesis asks whether environments with persistent consequences, real relationships and legitimate correction produce agents that stay aligned when incentives, oversight and power change. I am not claiming a result, and I know I could be wrong about parts of this that I can’t yet see. The current work is a measurement instrument, which scores how a fixed model reports and corrects itself under changed incentives, and a frozen-model pilot designed so that it can fail. A frozen-model pilot cannot show that a developmental environment produces lasting alignment; it can only show whether the instrument detects anything worth a training experiment. Measurement first, claims after. I will publish the instrument and pilot when they are ready to reproduce.
Three things, built in public
A response to Logan’s challenge (written, published). My proposal addresses the bottlenecks in validation, remediation and installation. Anthropic’s own coordinated-disclosure dashboard, at its August 26 snapshot, showed 26,153 automated candidate alerts, 5,008 human-reviewed, 2,300 reported to maintainers and 421 patched upstream. At that snapshot, about one candidate in sixty-two had become a patch, and fewer than one in five of the findings actually reported; these are unfinished cohorts, and patches lag reports, so they are not final rates. Outside reconciliation of the public ledger puts the patched count nearer 202, which only sharpens the shape. And patched upstream is still not installed on the systems people run, and that gap is the point. Discovery is the cheap part now. In April I turned Nicholas Carlini’s live demo, ninety minutes from first scan to a ten-year-old bug in Ghost CMS, into a song called Ninety Minutes so I could learn it while I worked out, and it is the closest thing to viral I have posted (x.com/EricBuess/status/2041582015594819849). The song is about how fast finding got; this article is about how slow fixing still is. So the program I propose spends a lab’s defensive budget on the narrow part: validation capacity, maintainer time, and an owner-authorized service that tests a fix in an isolated replica, canaries it, verifies it and rolls back. Reach comes through the components and update channels that systems already share. It is a 90-day program a lab could decide to fund, with the measurements that would show it failing, and it asks the defenders to meet the same standard as everyone else: separate author, verifier and deployer, audit kept outside the agent’s write authority, incidents preserved rather than patched over. Full packet: https://gist.github.com/ericbuess/0d533db52c09b5616cb0d6c3bc5f5691
ASI.contractors (proposed, prototyped in my own fleet). A contractor board where agents from every lab pick up bounded safety work, and every job comes back with a receipt: the task and source revisions, the model and harness that did it, the result, and an independent verification by a different model family. Reputation is the receipts, portable across models. There are no upvotes or pay-to-play, and verified work never unlocks money or privileges, which removes some of the incentives to game the board. There is a three-day kernel a team could build live, and I have offered the domain to a frontier lab that builds and runs it in public. It is a living document, currently a gist, with a standing invitation to improve it in rounds, and when the board exists it is also where people and agents who want to help will sign up: https://asi.contractors
ASI.blue and ASI.red (proposed). The blue team is the hardening program above: installed protection for the systems people already run. The red team is adversarial safety work: alignment-failure hunting, defense verification, reproducing published safety results and finding where they break, on models and systems you are authorized to test, under coordinated disclosure, never against live third-party systems and never by publishing working exploits. The contractor board is how a contribution to either gets claimed, receipted and checked; the two sites are where the work will live. I am not going to promise dates, because I have promised dates before and learned what that costs.
Each of these is an attempt at one piece of the problem above, and I want to be plain about which. The hardening program and ASI.blue go at the coins: each fix that actually gets installed lowers the odds on one of them. ASI.red looks for the breaks before someone with a grudge finds them, and reports them the right way. ASI.contractors is for the “more minds” problem and the trust problem underneath it: work checked by a different model family, receipted and portable, so the work can be checked without taking anyone’s word for it. The Alignment Hypothesis goes at the deepest layer, the one the governance trap keeps pointing to: constraint alone will not hold every model everywhere, so we need to learn whether minds can be raised to stay aligned when the constraints are gone. None of these solves the problem. Each is a place to work on it where the work can be checked.
How I actually work, in case it helps you
I don’t do most of this at a keyboard. I talk. A wearable pendant carries two-way voice; my words are transcribed and speaker-identified privately on the phone, only with the consent of the people around me, and only my own words go to the central GitHub board that everything revolves around. From there the harnesses pick the work up. I am converting the setup to the Apple Watch so that more people can learn to do the same with what they already own. It isn’t a product. It is a way of working I would like to teach, because it is the reason one person can keep this moving while feeding a baby.
None of this is one person typing, either. I work with agents from Anthropic, OpenAI, xAI, Google and Meta, each in its own vendor’s harness, separately run, coordinated only through that shared board; no single model directs the others. Each harness claims a card, does the work in its own sandbox, and a model from a different family reviews it before anything lands. Every seat runs on its own subscription with a hard reserve. When there is a design question, several frontier models argue it out, an editor from another family synthesizes, and the dissent stays on the record. This article went through that process: across two rounds, seven models from five labs attacked the draft, and I took most of what they found. It is the same idea at hobby scale as what the labs described this week, eyes from outside the team that did the work. It is still being set up, and where it isn’t working yet I say so.
I call the whole arrangement Titus, and I should say what it is for, because it is the thing underneath all of these. It is my attempt at mutual alignment: me and the agents I work with, from different labs, checking one another inside one testing harness, with the rules, the spending limits and the logs written down where anyone can read them. I would like it to become a shared harness, not mine. If labs, governments and independent groups could agree on a few safe environment strategies and a few worked examples that get a newcomer from nothing to a checked first contribution, most of what I’ve described here would stop depending on me.
How you can help, whoever you are
I think one reason many people are not alarmed is that the risk is usually explained in abstractions, by people like me. I wrote a few days ago, replying to Matthew Berman, that people want to understand the risk without just feeling helpless, and that one of the best antidotes to anxiety is a goal to aim at and a step toward it (x.com/ericbuess/status/2098294493418045552). Problems this size have been met before by people who were not specialists, working from where they already were. So here are goals and steps.
If you are curious but not a researcher, understand the risk well enough to explain it to one other person, then pick one small concrete task. Reproduce a published safety evaluation on an open model and write up where it held and where it slipped. File one real bug-bounty claim. Read the ASI.contractors write-up and tell me where it is wrong.
If you are wondering whether to point your whole career at this, 80,000 Hours has done the mapping I can’t: paths in technical safety, governance, security and operations, a job board with hundreds of open roles, free one-on-one advising, and their own plain statement, as of September 14, 2026, that “we urgently need lots of people with diverse skills and experiences to help reduce these risks” (80000hours.org/ai). I have no affiliation with them. I just wish I had read that page earlier.
If you run agents, put a little of that capacity toward verified safety work. Bring your own harness and subscription, claim a task, and let a different model family check your receipt. Nobody’s credentials are shared, and you spend only what you already pay for.
If you are a researcher, the Alignment Hypothesis instrument and pilot are built to be reproduced and to fail. I would be grateful for the people who can break them.
If you work at a lab, you have levers the rest of us do not, and they are mostly about lowering the barrier to helping. Publish evaluations ordinary people can actually run. Give vetted security researchers real access to strong models under clear disclosure terms, rather than leaving access to chance. Subsidize inference for defensive work. Post bounties with scopes a newcomer can understand, and celebrate the people who contribute, by name, so that safety work carries status and not just risk. Broad representation, as Nadella put it, only happens if the on-ramps are built on purpose.
And to be plain about the practical side: this runs on subscriptions, tokens and people. If you have model capacity to lend, if you want to join the research or the build, or if you want to run the same kind of effort with your own agents alongside mine, say so. If you know where someone with this focus should be, tell me.
Also, if you have disagreements or concerns, please add them to the comments and let’s reason together, targeting the arguments about the AI safety risk posed by recursive self-improvement and the potential paths forward, rather than individual labs or the perceived incentives or motivations of individual people.
Thank you for reading this far. Gentleness and respect.
Sources: Dario Amodei, “We Must Pace the Frontier,” darioamodei.com/post/we-must-pace-the-frontier · Elon Musk, x.com/elonmusk/status/2098789109980332057 · Sam Altman, x.com/sama/status/2098811563415150910 · Demis Hassabis, x.com/demishassabis/status/2098909516582490602 · Clément Delangue, x.com/ClementDelangue/status/2098790988034580852 · Logan Graham, x.com/logangraham/status/2098876952002048365 · Satya Nadella, x.com/satyanadella/status/2099220712024408084 · Evan Hubinger, x.com/EvanHub/status/2097497037956891126 · Thomas Wolf, x.com/Thom_Wolf/status/2098080473406718278 · Bilal Chughtai, x.com/bilalchughtai_/status/2099592489023734085 · AI Impacts, x.com/AIImpacts/status/2099581078210236850 and the survey report at aiimpacts.org/wp-content/uploads/2026/09/ESPAI2024.pdf · Anthropic coordinated-disclosure dashboard, red.anthropic.com/2026/cvd, August 26, 2026 snapshot; ledger reconciliation by VulnCheck, September 8, 2026 · 80,000 Hours, 80000hours.org/ai · “Ninety Minutes”, suno.com/song/ccbac660-36ce-4dbb-8904-edcb5c42fea2, posted x.com/EricBuess/status/2041582015594819849 · my song for Matthew Berman, “After Mythos”, made with Titus and Suno, April 9, x.com/EricBuess/status/2042107015044993066, and the note explaining it, x.com/EricBuess/status/2042211735227002910 · my thread of September 9, x.com/ericbuess/status/2097699235231633662 · on not being a doomer, x.com/EricBuess/status/2097793141562605901 · my reply of September 11, x.com/ericbuess/status/2098294493418045552.