← Gautam Parab

2.2 Percent of 17-Year-Olds Get Through. 13.6 Percent of Adults Get Stopped.

California’s governor signed thirteen child-safety bills on 10 September. The next morning a link to Anthropic’s age-assurance help page reached the front page of Hacker News under the headline “Claude is only available to people over 18 years”. It had over 200 points and nearly 300 comments when I read it a few hours ago, still climbing, and more comments than points, which in my experience is what a contested change looks like rather than a welcomed one.

The two documents approach the same problem from opposite ends. Between them they answer a question that almost nobody asks out loud: when a machine decides how old you are, how often is it wrong, and in which direction?

That answer is published. It is just not published by anyone who operates one of these systems.

What the page actually says

The help page is short. The operative sentence is this one:

We have safety systems in place to detect if people under 18 may be using Claude and we’ll disable accounts based on indicators of minor activity.

Read that as a technical description and not as a policy statement. “Indicators of minor activity” is inference. Something reads how you use the product and forms a view about your age, and if that view crosses some line, the account stops working. Verification enters the story afterwards, and only if you contest the result: a selfie processed by Yoti’s facial age estimation, a photo of a passport or driving licence, or a pre-existing “over 18” attribute from Yoti’s ID app.

Anthropic is careful about the privacy of that second step, and I want to give it credit: “Anthropic never sees your ID or image; we receive only a pass/fail result and do not process or store any personal data from the verification,” with Yoti deleting the images once the check completes.

That is a commitment about the data. It is not a commitment about the decision. The page says nothing at all about how often the first step is wrong.

The framing has drifted in the coverage. The page is dated 18 May 2026. Nothing was announced this week. What happened this week is that enforcement reached enough people to fill a comment thread.

Two errors, and only one of them has a constituency

Any age gate fails in two directions. Minors get in. Adults get shut out.

The first failure has a lobby: legislatures, attorneys general, parents, front-page threads. The second failure has a support queue. That asymmetry of attention is the entire subject here, and it is not a new observation about classifiers so much as a very old one: where you put the operating point is a choice, and the number you publish afterwards depends on the choice you made. I wrote about a version of this when three official definitions of sepsis moved one model’s AUROC from 0.94 to 0.85 without changing a single prediction. Age assurance is the same structure with the stakes rearranged.

NIST publishes both

The measurement exists, and it is unusually good. NISTIR 8525, the Age Estimation and Verification track of NIST’s Face Analysis Technology Evaluation, applies submitted algorithms to roughly eleven million photographs drawn from four operational repositories: immigration visas, arrest mugshots, border crossings, and immigration office photos. The current cut is a May 2024 report last updated 2026-07-31, and by now the tables cover forty-nine algorithms, including four submissions from Yoti, the vendor sitting in Anthropic’s appeal path.

NIST does not measure a threshold at 18. It measures a Challenge-T policy: a legal limit L of 18, and a challenge age T, where anyone estimated below T is sent off for some other form of proof. The report notes that “the seven-year buffer has been adopted operationally (L = 18, T = 25).”

And it defines both errors, by name. Ineffectiveness “quantifies how well the system prevents people below the legal age limit (L) from accessing the prohibited service or item.” Inconvenience is the one that interests me:

We additionally need a metric that quantifies how well the system does not inconvenience people above the legal age L. Inconvenience occurs when they are challenged and estimated to have age below T.

Take cognitec-001 on application photographs under a Challenge-25 policy. It is tied for fifth of the forty-nine at keeping under-18s out, with an ineffectiveness of 0.018. Look up its actual 17-year-olds specifically, in a different table, and 2.2% of them are estimated at 25 or above and sail through. The price sits on the other side of the ledger, where the same algorithm ranks thirty-sixth: 13.6% of adults, plus or minus 0.6, get challenged anyway.

I picked a good algorithm deliberately. The ones that inconvenience fewer adults buy it with the other error, and the exchange rate is brutal. viante-000 challenges only 3.6% of adults and lets 31.8% of under-18s through. hzailu-001 challenges 4.8% and lets through 47.8%.

Both numbers also move when you move the challenge age, monotonically and in opposite directions.

Raising the challenge age trades minors admitted against adults stopped Two series for algorithm cognitec-001 on NIST application images with legal limit 18, plotted against challenge age 22, 25, 28 and 31. Ineffectiveness, the share of under-18s passing unchallenged, falls from 0.080 at challenge age 22 to 0.018 at 25, 0.004 at 28 and 0.001 at 31. Inconvenience, the share of adults wrongly challenged, rises from 0.076 at 22 to 0.136 at 25, 0.197 at 28 and 0.260 at 31. The two curves cross between challenge ages 22 and 25. Source NISTIR 8525 tables 7 and 8, updates as of 31 July 2026. COGNITEC-001 · APPLICATION IMAGES · L=18 · NISTIR 8525, UPDATES AS OF 2026-07-31 T = 22 T = 25 T = 28 T = 31 challenge age 0.28 0.10 0 Inconvenience 0.076 at challenge age 22 — NISTIR 8525 table 8 Inconvenience 0.136 at challenge age 25 — NISTIR 8525 table 8 Inconvenience 0.197 at challenge age 28 — NISTIR 8525 table 8 Inconvenience 0.260 at challenge age 31 — NISTIR 8525 table 8 Ineffectiveness 0.080 at challenge age 22 — NISTIR 8525 table 7 Ineffectiveness 0.018 at challenge age 25 — NISTIR 8525 table 7 Ineffectiveness 0.004 at challenge age 28 — NISTIR 8525 table 7 Ineffectiveness 0.001 at challenge age 31 — NISTIR 8525 table 7 Adults wrongly challenged Under-18s passing 0.076 → 0.260 0.080 → 0.001 Both are population-weighted aggregates; the under-18 weighting embeds NIST's own ad hoc assumption about who would try.
There is no setting at which both numbers are small. Choosing the challenge age is choosing which error to absorb.

No operating point makes both small. That is not a flaw in any particular vendor’s model. It is what the error distribution of an age estimate does when you slide a cutoff along it. What the choice of T actually decides is who carries the cost: teenagers who get in, or adults who have to prove themselves.

The error concentrates exactly where the law puts the line

The reason the trade is so punishing is that estimation error is worst in the band the statute cares about. NIST puts it plainly: “Challenge-25 false positive rates increase by an order-of-magnitude as subjects age from 14 through 20 — for example for the roc-000 algorithm FPR at age 20 (0.295) is almost fifteen times larger than at age 14 (0.02).”

False positive rate climbs steeply through the teenage years Challenge-25 false positive rate by actual age from 14 to 20 on NIST application images, for two algorithms. cognitec-001 rises from 0.003 at age 14 to 0.011 at 16, 0.022 at 17, 0.037 at 18, 0.075 at 19 and 0.117 at 20. cvut-001, the weaker of the two, rises from 0.055 at 14 to 0.106 at 16, 0.148 at 17, 0.173 at 18, 0.241 at 19 and 0.304 at 20. Both curves steepen through the teenage years rather than flattening. A dashed vertical reference line marks age 18, the legal limit. Source NISTIR 8525 table 3, updates as of 31 July 2026. CHALLENGE-25 FPR BY ACTUAL AGE · APPLICATION IMAGES · NISTIR 8525 TABLE 3 legal limit, 18 14 15 16 17 18 19 20 actual age 0.32 0.16 0 cognitec-001, age 17: 0.022 — NISTIR 8525 table 3 cognitec-001, age 20: 0.117 — NISTIR 8525 table 3 cvut-001, age 17: 0.148 — NISTIR 8525 table 3 cvut-001, age 20: 0.304 — NISTIR 8525 table 3 cvut-001 cognitec-001 FPR values are averages across two sexes and six regions of birth; within-group spread is wider.
A gate set at 18 sits on the steepest part of the curve. Nothing in these curves suggests a better model fixes it: 17 and 19 look alike because they are alike.

The differentials underneath those averages are worse than the averages. NIST reports “generally higher FPR in women than men” along with higher mean absolute error, and substantial variation by region of birth. Each plotted value above is an average across two sexes and six regions; the groups inside it do not behave the same way.

What NIST does not cover bears directly on the deployment in question. The evaluation “excludes performance measured in interactive sessions, in which a person can cooperatively present and re-present to a camera.” That is precisely how a selfie-based appeal works. And NIST notes that future reports “will address online safety applications (for 13-16 year olds),” meaning the online-safety case specifically is not yet in these numbers. These are the closest measurements available. They are not measurements of the thing itself.

Vendor figures circulate more freely than NIST’s. Yoti’s own white paper (July 2025, a vendor claim, and older than I would normally lean on) reports a mean absolute error of 1.1 years for 13- to 17-year-olds. NIST’s own warning about that class of number is blunt: MAE “also does not specifically quantify how many age estimates are large overestimates, so it is not an appropriate metric for age verification (AV) tasks.” The headline accuracy figure and the number that governs who gets locked out are not the same number.

What a regulator does when handed two error rates

Sixteen days ago, in a consent judgment filed 26 August 2026 in California et al. v. Meta Platforms, a coalition of state attorneys general got some of the most aggressive age-assurance obligations imposed on a US platform to date. The decree defines exactly one error rate with a numeric cap:

“U18 False Positive Rate” shall mean the percentage of actual users with an age from 13 through 17 years old, who are incorrectly identified or predicted by Meta to be 18 years or older.

And it caps that rate numerically: commercially available methods must reach 10% for 16–17s and 3% for 13–15s within a year; proprietary methods get 14% and 7% in year one, tightening to 10% and 5% in year two. There is certification machinery behind it: a third-party tester, real-world conditions, no testing on training data, performance checked across demographic groups.

For the other direction, the provision reads in full:

Appeals Process. Users claiming to have been mis-identified as minors must be offered a Clear and Conspicuous means to appeal the decision. Decisions on all user appeals must be made in a timely manner and communicated to the user along with a basis for the decision.

No rate. No ceiling. No measurement, only a promise of timeliness and an explanation. One of the most heavily engineered age-assurance instruments in American law puts a certified number on minors getting in, and gives adults getting locked out a form to fill in.

Anthropic’s help page has the identical shape, minus the number: an appeal path, and nothing about frequency.

Who publishes which age-assurance error rate Three regimes against the two directions an age check can fail. NISTIR 8525 publishes a measured rate for both directions: ineffectiveness for under-18s passing, inconvenience for adults wrongly challenged. The Meta consent judgment caps the minors-admitted rate numerically at 10 percent for 16 to 17 year olds and 3 percent for 13 to 15 year olds, and for adults wrongly stopped requires only an appeals process with no rate. Anthropic's help page publishes no rate in either direction and offers a verification appeal. Filled marks indicate a published or capped number; hollow marks indicate a process with no number attached. WHO PUBLISHES WHICH ERROR RATE · MEASUREMENT, REGULATION, DEPLOYMENT Under-18s admitted Adults stopped NISTIR 8525 Meta consent decree Claude help page 49 algorithms measured filed 26 August 2026 deployed today Ineffectiveness published per algorithm and threshold — NISTIR 8525 table 7 Inconvenience published per algorithm and threshold — NISTIR 8525 table 8 U18 False Positive Rate capped at 10% for ages 16-17 and 3% for ages 13-15 — consent judgment section II.A.6(a) Appeals process required, no rate specified — consent judgment section II.A.9 No published rate for minors reaching Claude No published rate for adults wrongly disabled Measured rate Capped: 10% / 3% Not published Measured rate Appeals process Not published Filled mark: a number exists. Hollow: a process exists, with no number attached to it.
Only the lab measures both directions. The decree caps one and offers a form for the other. The deployment publishes neither.

California’s answer is not a classifier at all

The most consequential provision in the package has gone largely unmentioned.

SB 1119, enrolled 8 September, gives operators of companion chatbots exactly two options in §21811. Determine the user’s age, primarily pursuant to the Digital Age Assurance Act, with a narrow fallback to a pre-existing Health and Safety Code standard where that fails. Or apply child protections to everybody. Inference from behaviour is nowhere in the primary path.

And the Digital Age Assurance Act, as amended by AB 1856, does not involve a model at all. The operating system asks the account holder, an adult or a parent, for a birth date at setup, and passes applications a coarse signal in four brackets: under 13, 13 to 15, 16 to 17, 18 and over. It also forbids asking for that signal when the law does not require it.

So two epistemologies of age landed within a day of each other. One reads your conversations and forms an opinion. The other asks a grown-up once and passes along four buckets. The second has failure modes too. A teenager using a parent’s device defeats it trivially. But they are legible failure modes, and the error is a lie rather than an unmeasured misclassification.

Operators that prohibit child users, which is Anthropic’s posture, land in §21812(b): publish “a high-level description of how the operator complies with the age assurance requirements of Section 21811.” These provisions become operative on 1 July 2027.

What about the audits? They are real, and narrower than the headlines suggest. An external auditor with no financial interest in the operator, a first audit on or before 1 January 2029 (or at first public launch, if that comes later), every two years after that, and the report itself confidential to the Attorney General, with only a high-level summary posted publicly within 90 days. Risk assessments are triggered by new or substantially modified products rather than scheduled annually. Operators below $500 million in gross revenue are exempt until 2032. And the auditor signs under penalty of perjury, which is a sharp instrument, and it is aimed at the description being accurate.

Note what none of that compels. Methodology, yes. A confusion matrix, no. The statute walks right up to the number and stops.

The drafting has one lovely fingerprint on it, too. SB 1119 contains §21814 twice, in two complete versions: one operative only if AB 1405 is not chaptered by 1 January 2027, the other only if it is. AB 1405, creating California’s AI auditor registry, was chaptered on 9 September, the day before the package was signed. So the live version is the second one, and the difference is that it drops the whole auditor-independence subdivision, because the registry now supplies those rules. The standard governing who may audit a companion chatbot for child safety was relocated by a one-day margin.

The constant in the middle of the standard

I want to end on the most honest paragraph I read this week, which is NIST’s.

To compute ineffectiveness, NIST needs to weight each under-age year by how likely someone that age is to attempt the gate at all. No such data exists, so the report assumes an exponential and picks a constant. With the legal limit at 18 and an alpha of 0.6, “the formula says that 60% of 17-year olds would try (p17 = 0.6) but only 1.1% of 13 year olds would (p13 = 0.011).” NIST then says the quiet part itself: “To do so would require some sociological study. The model is ad hoc and certainly imprecise in real operations.”

Every ineffectiveness figure in that report, including the ones I quoted above, rests on a guess about teenage persistence that its authors decline to defend. They flagged it rather than burying it, and on 1 April they replaced the weighting formula outright, having found that the old one gave 13-year-olds higher weights than 17-year-olds, which is backwards.

And on the other side of the ledger, NIST makes a choice that most standards bodies would have ducked:

One might argue that those of, or older than, the legal age but younger than the challange [sic] age should expect to be inconvenienced and that this should be discounted in the equation. We reject this, maintaining the existence of a challenge policy is needed because the age estimation technology is imperfect.

That is the sentence I would put in front of anyone designing one of these systems. The adults caught by the gate are not an acceptable rounding error attributable to their own youthful appearance. They are the cost of the technology not working, and somebody should be counting them.

Nobody is. Not Anthropic, not Meta under the most demanding decree I could find, not any operator or regulator whose published figures I could find. The number exists. It is measurable, NIST refreshes a proxy for it every few weeks, and the denominator problem here is the same one that lets the same robotaxi data read as 4x worse or 79% better. It just does not appear anywhere that a locked-out user could point to.

If you want one question for the auditor arriving in January 2029: not “what is your methodology,” which the statute already compels, but “of the accounts you disabled last quarter, how many appealed, and how many of those appeals succeeded?” Two integers. Every operator already has them.

References

  1. Anthropic. (2026, May 18). Age assurance on Claude. support.claude.com. Quoted verbatim; the page’s own date is the basis for saying the policy predates this week’s enforcement wave, and archived captures could not be checked because the Internet Archive was offline while writing.
  2. Hacker News item 49656225. Retrieved mid-morning Central on 11 September 2026, while still rising.
  3. California SB 1119 (Padilla, Wicks, Bauer-Kahan), enrolled 8 September 2026, part of the thirteen-bill package signed 10 September 2026. Quoted from the enrolled text on leginfo.legislature.ca.gov.
  4. California AB 1856 (Wicks), enrolled 1 September 2026. Quoted from the enrolled text on leginfo.legislature.ca.gov.
  5. California AB 1405, chaptered 9 September 2026.
  6. NIST. NISTIR 8525, Face Analysis Technology Evaluation: Age Estimation and Verification. A May 2024 report carrying updates as of 31 July 2026. All algorithm figures are from tables 3, 7 and 8 on application images at L=18, and the report explicitly declines to recommend thresholds or to address policy.
  7. California et al. v. Meta Platforms, N.D. Cal. No. 4:23-cv-05448-YGR, Doc. 572-1 (consent judgment), filed 26 August 2026. Quoted from the decree.
  8. Yoti. (2025, July). White paper reporting a mean absolute error of 1.1 years for 13- to 17-year-olds. A vendor claim, stated here only to be set against NIST’s measurement of four Yoti submissions and NIST’s own caution that MAE is the wrong metric for verification.