California’s governor signed thirteen child-safety bills on 10 September. The next morning a link to Anthropic’s age-assurance help page reached the front page of Hacker News under the headline “Claude is only available to people over 18 years”. It had over 200 points and nearly 300 comments when I read it a few hours ago, still climbing, and more comments than points, which in my experience is what a contested change looks like rather than a welcomed one.
The two documents approach the same problem from opposite ends. Between them they answer a question that almost nobody asks out loud: when a machine decides how old you are, how often is it wrong, and in which direction?
That answer is published. It is just not published by anyone who operates one of these systems.
The help page is short. The operative sentence is this one:
We have safety systems in place to detect if people under 18 may be using Claude and we’ll disable accounts based on indicators of minor activity.
Read that as a technical description and not as a policy statement. “Indicators of minor activity” is inference. Something reads how you use the product and forms a view about your age, and if that view crosses some line, the account stops working. Verification enters the story afterwards, and only if you contest the result: a selfie processed by Yoti’s facial age estimation, a photo of a passport or driving licence, or a pre-existing “over 18” attribute from Yoti’s ID app.
Anthropic is careful about the privacy of that second step, and I want to give it credit: “Anthropic never sees your ID or image; we receive only a pass/fail result and do not process or store any personal data from the verification,” with Yoti deleting the images once the check completes.
That is a commitment about the data. It is not a commitment about the decision. The page says nothing at all about how often the first step is wrong.
The framing has drifted in the coverage. The page is dated 18 May 2026. Nothing was announced this week. What happened this week is that enforcement reached enough people to fill a comment thread.
Any age gate fails in two directions. Minors get in. Adults get shut out.
The first failure has a lobby: legislatures, attorneys general, parents, front-page threads. The second failure has a support queue. That asymmetry of attention is the entire subject here, and it is not a new observation about classifiers so much as a very old one: where you put the operating point is a choice, and the number you publish afterwards depends on the choice you made. I wrote about a version of this when three official definitions of sepsis moved one model’s AUROC from 0.94 to 0.85 without changing a single prediction. Age assurance is the same structure with the stakes rearranged.
The measurement exists, and it is unusually good. NISTIR 8525, the Age Estimation and Verification track of NIST’s Face Analysis Technology Evaluation, applies submitted algorithms to roughly eleven million photographs drawn from four operational repositories: immigration visas, arrest mugshots, border crossings, and immigration office photos. The current cut is a May 2024 report last updated 2026-07-31, and by now the tables cover forty-nine algorithms, including four submissions from Yoti, the vendor sitting in Anthropic’s appeal path.
NIST does not measure a threshold at 18. It measures a Challenge-T policy: a legal limit L of 18, and a challenge age T, where anyone estimated below T is sent off for some other form of proof. The report notes that “the seven-year buffer has been adopted operationally (L = 18, T = 25).”
And it defines both errors, by name. Ineffectiveness “quantifies how well the system prevents people below the legal age limit (L) from accessing the prohibited service or item.” Inconvenience is the one that interests me:
We additionally need a metric that quantifies how well the system does not inconvenience people above the legal age L. Inconvenience occurs when they are challenged and estimated to have age below T.
Take cognitec-001 on application photographs under a Challenge-25 policy. It is tied for fifth of the forty-nine at keeping under-18s out, with an ineffectiveness of 0.018. Look up its actual 17-year-olds specifically, in a different table, and 2.2% of them are estimated at 25 or above and sail through. The price sits on the other side of the ledger, where the same algorithm ranks thirty-sixth: 13.6% of adults, plus or minus 0.6, get challenged anyway.
I picked a good algorithm deliberately. The ones that inconvenience fewer adults buy it with the other error, and the exchange rate is brutal. viante-000 challenges only 3.6% of adults and lets 31.8% of under-18s through. hzailu-001 challenges 4.8% and lets through 47.8%.
Both numbers also move when you move the challenge age, monotonically and in opposite directions.
No operating point makes both small. That is not a flaw in any particular vendor’s model. It is what the error distribution of an age estimate does when you slide a cutoff along it. What the choice of T actually decides is who carries the cost: teenagers who get in, or adults who have to prove themselves.
The reason the trade is so punishing is that estimation error is worst in the band the statute cares about. NIST puts it plainly: “Challenge-25 false positive rates increase by an order-of-magnitude as subjects age from 14 through 20 — for example for the roc-000 algorithm FPR at age 20 (0.295) is almost fifteen times larger than at age 14 (0.02).”
The differentials underneath those averages are worse than the averages. NIST reports “generally higher FPR in women than men” along with higher mean absolute error, and substantial variation by region of birth. Each plotted value above is an average across two sexes and six regions; the groups inside it do not behave the same way.
What NIST does not cover bears directly on the deployment in question. The evaluation “excludes performance measured in interactive sessions, in which a person can cooperatively present and re-present to a camera.” That is precisely how a selfie-based appeal works. And NIST notes that future reports “will address online safety applications (for 13-16 year olds),” meaning the online-safety case specifically is not yet in these numbers. These are the closest measurements available. They are not measurements of the thing itself.
Vendor figures circulate more freely than NIST’s. Yoti’s own white paper (July 2025, a vendor claim, and older than I would normally lean on) reports a mean absolute error of 1.1 years for 13- to 17-year-olds. NIST’s own warning about that class of number is blunt: MAE “also does not specifically quantify how many age estimates are large overestimates, so it is not an appropriate metric for age verification (AV) tasks.” The headline accuracy figure and the number that governs who gets locked out are not the same number.
Sixteen days ago, in a consent judgment filed 26 August 2026 in California et al. v. Meta Platforms, a coalition of state attorneys general got some of the most aggressive age-assurance obligations imposed on a US platform to date. The decree defines exactly one error rate with a numeric cap:
“U18 False Positive Rate” shall mean the percentage of actual users with an age from 13 through 17 years old, who are incorrectly identified or predicted by Meta to be 18 years or older.
And it caps that rate numerically: commercially available methods must reach 10% for 16–17s and 3% for 13–15s within a year; proprietary methods get 14% and 7% in year one, tightening to 10% and 5% in year two. There is certification machinery behind it: a third-party tester, real-world conditions, no testing on training data, performance checked across demographic groups.
For the other direction, the provision reads in full:
Appeals Process. Users claiming to have been mis-identified as minors must be offered a Clear and Conspicuous means to appeal the decision. Decisions on all user appeals must be made in a timely manner and communicated to the user along with a basis for the decision.
No rate. No ceiling. No measurement, only a promise of timeliness and an explanation. One of the most heavily engineered age-assurance instruments in American law puts a certified number on minors getting in, and gives adults getting locked out a form to fill in.
Anthropic’s help page has the identical shape, minus the number: an appeal path, and nothing about frequency.
The most consequential provision in the package has gone largely unmentioned.
SB 1119, enrolled 8 September, gives operators of companion chatbots exactly two options in §21811. Determine the user’s age, primarily pursuant to the Digital Age Assurance Act, with a narrow fallback to a pre-existing Health and Safety Code standard where that fails. Or apply child protections to everybody. Inference from behaviour is nowhere in the primary path.
And the Digital Age Assurance Act, as amended by AB 1856, does not involve a model at all. The operating system asks the account holder, an adult or a parent, for a birth date at setup, and passes applications a coarse signal in four brackets: under 13, 13 to 15, 16 to 17, 18 and over. It also forbids asking for that signal when the law does not require it.
So two epistemologies of age landed within a day of each other. One reads your conversations and forms an opinion. The other asks a grown-up once and passes along four buckets. The second has failure modes too. A teenager using a parent’s device defeats it trivially. But they are legible failure modes, and the error is a lie rather than an unmeasured misclassification.
Operators that prohibit child users, which is Anthropic’s posture, land in §21812(b): publish “a high-level description of how the operator complies with the age assurance requirements of Section 21811.” These provisions become operative on 1 July 2027.
What about the audits? They are real, and narrower than the headlines suggest. An external auditor with no financial interest in the operator, a first audit on or before 1 January 2029 (or at first public launch, if that comes later), every two years after that, and the report itself confidential to the Attorney General, with only a high-level summary posted publicly within 90 days. Risk assessments are triggered by new or substantially modified products rather than scheduled annually. Operators below $500 million in gross revenue are exempt until 2032. And the auditor signs under penalty of perjury, which is a sharp instrument, and it is aimed at the description being accurate.
Note what none of that compels. Methodology, yes. A confusion matrix, no. The statute walks right up to the number and stops.
The drafting has one lovely fingerprint on it, too. SB 1119 contains §21814 twice, in two complete versions: one operative only if AB 1405 is not chaptered by 1 January 2027, the other only if it is. AB 1405, creating California’s AI auditor registry, was chaptered on 9 September, the day before the package was signed. So the live version is the second one, and the difference is that it drops the whole auditor-independence subdivision, because the registry now supplies those rules. The standard governing who may audit a companion chatbot for child safety was relocated by a one-day margin.
I want to end on the most honest paragraph I read this week, which is NIST’s.
To compute ineffectiveness, NIST needs to weight each under-age year by how likely someone that age is to attempt the gate at all. No such data exists, so the report assumes an exponential and picks a constant. With the legal limit at 18 and an alpha of 0.6, “the formula says that 60% of 17-year olds would try (p17 = 0.6) but only 1.1% of 13 year olds would (p13 = 0.011).” NIST then says the quiet part itself: “To do so would require some sociological study. The model is ad hoc and certainly imprecise in real operations.”
Every ineffectiveness figure in that report, including the ones I quoted above, rests on a guess about teenage persistence that its authors decline to defend. They flagged it rather than burying it, and on 1 April they replaced the weighting formula outright, having found that the old one gave 13-year-olds higher weights than 17-year-olds, which is backwards.
And on the other side of the ledger, NIST makes a choice that most standards bodies would have ducked:
One might argue that those of, or older than, the legal age but younger than the challange [sic] age should expect to be inconvenienced and that this should be discounted in the equation. We reject this, maintaining the existence of a challenge policy is needed because the age estimation technology is imperfect.
That is the sentence I would put in front of anyone designing one of these systems. The adults caught by the gate are not an acceptable rounding error attributable to their own youthful appearance. They are the cost of the technology not working, and somebody should be counting them.
Nobody is. Not Anthropic, not Meta under the most demanding decree I could find, not any operator or regulator whose published figures I could find. The number exists. It is measurable, NIST refreshes a proxy for it every few weeks, and the denominator problem here is the same one that lets the same robotaxi data read as 4x worse or 79% better. It just does not appear anywhere that a locked-out user could point to.
If you want one question for the auditor arriving in January 2029: not “what is your methodology,” which the statute already compels, but “of the accounts you disabled last quarter, how many appealed, and how many of those appeals succeeded?” Two integers. Every operator already has them.
References