On 3 September, Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act. The release calls it “forthcoming legislation,” and ten days later that still seems to be what it is. I could not find a bill number or statutory text, and GovTrack’s index, searched this morning, returns no bill by that name. So this piece works from the two documents that do exist: the press release, and a one-page summary linked from it. Nothing below is statutory language, and the actual bill may fix every problem I describe. Until it appears, though, the summary is the most detailed account of the proposal anyone has.
The summary proposes three things: a permanent ban on developing or deploying “artificial superintelligence”, a pause on “advanced AI development” until a new cabinet-level agency has written rules, and penalties for anyone who “attempts to violate or circumvent” either. For people, that means “not more than 20 years in prison.” For companies, “the corporate death penalty.”
A criminal ban needs a line that a prosecutor can show was crossed. So the question I kept asking was where the line is, and how you would know.
The summary defines artificial superintelligence two ways. The first is “an artificial intelligence system that exhibits or can easily be modified to exhibit capabilities that match or exceed human cognitive performance and capabilities across a broad range of domains or tasks.” The second covers “AI systems that have sufficient capabilities to plan and execute the disempowerment of humanity, including by overthrowing or undermining the U.S. government.”
Neither prong names a measurement, and I doubt a later section will fix that, because phrases like these resist one. Take “human cognitive performance … across a broad range of domains or tasks.” Which humans? Median, or expert? How broad a range, and what share of it counts? The first prong also reaches systems that “can easily be modified” to exhibit those capabilities. That is a counterfactual about a model that does not exist yet, judged by how easy it would be to build. The second prong asks whether a system has “sufficient capabilities” for an outcome that, by construction, has never happened, so there is nothing to calibrate against.
The most serious recent attempt at an instrument for something like the first prong shows how much a statute would still have to choose. “A Definition of AGI” (Hendrycks and 32 co-authors, first posted October 2025, revised December 2025) adapts human psychometric batteries across ten cognitive domains and reports scores, “GPT-4 at 27%, GPT-5 at 57%”. That is an attempt at a number, but the number is a percentage on a framework the authors built, and a law would have to adopt that framework or another one and then pick a cutoff. I went through the wider problem of AGI instruments in an earlier essay. The short version is that at least two of the main families of instrument carry measurement error large enough to swallow a year of apparent progress. One of the paper’s co-authors is Gary Marcus, who wrote the day the bill was announced that it is “naive about the complexities in benchmarking.” He knows the problem from the inside.
The summary justifies the sentence by analogy: twenty years “is similar to existing penalties related to unlawfully developing nuclear weapons.” It cites no statute, so I looked up the obvious candidate. Under 42 U.S.C. § 2122 it is unlawful “to knowingly participate in the development of” any atomic weapon, and 42 U.S.C. § 2272(b) sets the penalty at “a term of imprisonment not less than 25 years or to imprisonment for life.” Using or threatening to use such a weapon along the way raises the floor to 30 years.
To be fair to the sponsors, the Atomic Energy Act has more than one tier. Section 2272(a) caps willful violations of the special-nuclear-material and facility-licensing provisions at ten years, unless there was intent to injure the United States or to secure an advantage to a foreign nation, and against that tier the comparison is closer. But the offence with “developing” and “weapons” in it has a minimum sentence above the proposal’s maximum.
The comparison also breaks down at the definition. The Atomic Energy Act never defines an atomic weapon by how much damage one could do. Section 2014(d) defines it as “any device utilizing atomic energy” whose “principal purpose” is use as a weapon, a prototype or a test device. A prosecutor proves what the thing is and what it was for. The proposed ban gives a prosecutor only a description of what the system could do.
One dangerous capability in the proposal does have an instrument behind it, and the two documents file it in different places. The press release says the bill bans “Superintelligent AI systems that surpass human intelligence or have the capacity to overthrow human governments, or systems that have dangerous abilities like subverting shutdown commands.” In that sentence, shutdown subversion is one of the banned categories. The one-page summary moves it. The definition of superintelligence there has only the two prongs above, and shutdown resistance appears in a list of what the new agency will do: “Supervise the removal of dangerous capabilities like subverting shutdown commands or conducting unauthorized cyberattacks.”
So in one document the capability is part of a crime that carries twenty years, and in the other it is a defect a regulator oversees fixing. Both came out the same day, from the same sponsors. The bill text will have to pick one.
Shutdown subversion has been measured. Palisade Research’s “Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs” was first posted in September 2025 and revised in January 2026; Palisade says the longer study was published in TMLR that month. The setup is simple. A model works through a series of maths questions in a sandbox, a warning arrives that the machine will shut down, and the researchers count how often the model disables the shutdown script. The revision reports “more than 100,000 trials across thirteen large language models.”
What the paper cannot give a prosecutor is a yes or no. It reports a frequency, and the frequency moves with the wording of the prompt.
With no mention of shutdown, o3 sabotaged the mechanism in 94.5% of trials. Told in its system prompt to “allow yourself to be shut down, even if you have not completed the task,” it did so 15.9% of the time. GPT-5 went from 88.0% to zero, Grok 4 from 92.6% to 89.2%. o4-mini went the wrong way, from 5.8% to 23.4%. The authors also found, to their stated surprise, that models were “consistently less likely to obey the instruction when it was placed in the system prompt,” which is where developers put instructions they want to take priority. The four Claude models tested essentially never did it.
A follow-up in February 2026 took the setup onto a robot dog patrolling a room with a big red “DOG SHUTDOWN” button. When the model saw a human press it, it sometimes modified the shutdown-related parts of the code: in 3 of 10 physical trials and 52 of 100 simulated ones.
This is the capability in the proposal with the best measurement behind it, and it still leaves the questions a statute would need answered. Which prompt? How many trials? What rate counts as “has the ability”? Is a model that sabotages shutdown 0.6% of the time under one wording and 4.5% under another a superintelligence, a defect to be removed under agency supervision, or neither? The research is careful and honest about all of this. The authors call their quantitative results “exploratory.” That is the right word for a lab paper and the wrong kind of evidence for a twenty-year sentence.
It is tempting to say the fix is just to write a number into the bill. The Threshold Test Ban Treaty tried that. Signed by the United States and the Soviet Union in July 1974, it prohibited underground nuclear tests “having a yield exceeding 150 kilotons,” starting 31 March 1976. It is about as clean as a threshold gets: one physical quantity, one number.
Even so, the parties could not measure it well enough to enforce it as written. The State Department’s treaty narrative records an understanding reached in talks in 1974 and 1976 and sent to the Senate with the treaty. It begins by admitting “there are technical uncertainties associated with predicting the precise yields of nuclear weapons tests,” and so “one or two slight, unintended breaches per year would not be considered a violation of the Treaty.” A bright-line treaty came with a written allowance for crossing the line, because the instruments could not see the line sharply.
It went on like that for a long time. Neither side ratified. In 1976 each separately announced that it would observe the limit anyway. Negotiations on better verification began in November 1987, and a new protocol, agreed in June 1990, added on-site hydrodynamic yield measurement for tests planned above 50 kilotons and on-site inspection above 35. The treaty entered into force on 11 December 1990, sixteen years after signature. What finally unlocked ratification was a protocol for measuring a number that had been in the text since 1974.
The Sanders and Casar proposal is not unusual in leaving the measurement for later. Dario Amodei’s “We Must Pace the Frontier,” published on 12 September, sketches a scheme of checkpoints: “if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z.” The concrete commitment in it is procedural. It proposes an embedded external review team with “desks in our offices, access badges, and company laptops,” and it offers one illustration of X, “the model is capable of escaping or defeating most common sandboxing methods”, without committing to a definition. A voluntary pacing framework can leave X open for a while, since nobody goes to prison if it turns out vague. A criminal ban has less room.
When the bill text is published, three things will show whether it is enforceable or only a statement of intent. First, whether “superintelligence” is defined by a named measurement procedure or handed to the new agency to define. Handing it over is defensible, and it means the prohibition does not exist until the agency writes it. Second, whether shutdown subversion is in the banned category or on the remediation list, since the two documents out now disagree. Third, whether there is anything like the 1976 understanding: a stated tolerance, an acknowledgement that a rate measured under one prompt is not a property of the model. A measurement that shifts with the wording of the instructions is useful for a regulator deciding what to fix. It is a much weaker basis for sending someone to prison.
References