The phrase that travelled was “abrupt losses in cognitive competence.” It comes from the abstract of “Large-Language Models as a Cognitive Virus,” posted to arXiv on 3 September 2026 by nine authors including David Krakauer, who runs the Santa Fe Institute, and Michael Levin at Tufts. Two days later it was on the front page of Hacker News, where it had collected 268 points and 196 comments when I looked. It is a good paper with a memorable title, and the memorable title is going to outrun the paper.
Here is what the model is. Three compartments: U, uncoupled users with little or no LLM use; C, coupled-autonomous users who use the tools regularly but retain independent cognition; and D, persistently dependent users for whom the interaction dominates. People move between them at rates — λ from U to C, ρ back again, μ from C down to D, σ from D back up to C — plus a fifth parameter, κ, weighting a nonlinear term that runs the other way. Each compartment carries a competence weight, and the paper sets those to 1.0, 0.5 and 0.1.
κ is worth pausing on, because it is the parameter that resists the headline. The term κU²C is added to the uncoupled population and subtracted from the coupled one, which makes it a second return flow alongside ρ, scaling with the square of how many uncoupled people there already are. The authors call it an Allee-like cooperative mechanism: independent cognitive practice is socially reinforced, so “cultural expectations that reward independent reasoning become more effective when autonomous individuals are common.” Raising κ makes adoption harder to establish, not easier.
That is the whole apparatus, and the apparatus is fine. Compartmental models are a respectable way to reason about how something spreads, and this one is put together by people who know what they are doing — Krakauer at Santa Fe, Levin at Tufts, Ricard Solé at ICREA and Pompeu Fabra, Manlio De Domenico at Padua. The paper is also honest about its own status. When it orders the compartments by competence it says plainly: “This ordering is an illustrative modelling assumption, not a general claim about LLM use.” It concedes that its mean-field maths sits on top of a network topology that real populations do not have.
The problem is the distance between that sentence and the sentence about abrupt losses in cognitive competence, and which of the two gets quoted.
Run down the five parameters and ask what it would take to put a number on any of them.
σ, the recovery rate from dependence back to autonomy, is the one I keep getting stuck on. To measure it you need to define a recovery event: a person who was persistently dependent and now is not. Over what interval? Dependent on what operational test — self-report, task performance without the tool, some behavioural trace? Measured against what baseline, given that the same person’s unaided competence was never recorded before they started? Epidemiology gets σ from people who were sick and then were not, and it knows which is which because there is a clinical definition of both states. There is no clinical definition of cognitive dependence on a chatbot, which means there is no denominator, which means there is no rate. It is the same hole I wrote about in what per-mile crash statistics can and cannot show: the number is unavailable because the category has not been operationalised.
μ has the same problem in the other direction. κ is worse — it asks how strongly the prevalence of autonomous people reinforces someone else’s return to autonomy, which would require watching the same population at different compositions and measuring how the return rate changes. λ is the only one with a plausible path to measurement, and even there the available data measures the wrong thing.
Pew put the stocks on the record in June: 49% of US adults say they have used an AI chatbot, 44% have used ChatGPT, 24% use one daily, 38% of employed adults use one for work, from a sample of 5,119 surveyed in February. Those are real measurements and they tell you nothing about transitions. A cross-section counts who is in each compartment on one day. The model runs on how fast people move between compartments, and nobody has followed a panel of users long enough to say.
It is not that the states are unmeasurable. Li and colleagues published a validated Large Language Model Dependence Scale in 2025, an eighteen-item bifactor instrument separating what they call functional dependence from existential dependence, developed on an exploratory sample of 421 users and a confirmatory sample of 1,030. That is a serviceable definition of compartment D. Point it at the same people every quarter for two years and you would have μ and σ. Nobody has, so the instrument measures a state and the model still runs on flows.
The deeper problem is that this class of model produces thresholds whether or not the world has any.
Bissell and colleagues showed the mechanism cleanly in 2014, in a compartmental model of smoking. They took two rates that had previously been constant — the rate at which smokers quit, and the rate at which former smokers relapse — and made them depend on peer influence. Their own summary of the result is the whole argument in one sentence: it “not only modifies the number of steady-states and nature of their asymptotic stability, but also introduces a new kind of non-linear ‘tipping-point’ dynamic.” The tipping point arrived with the functional form, not with any new data.
Iacopini and colleagues made the same point from a different angle in Nature Communications in 2019: hold the parameters identical and switch the interactions from pairwise to higher-order, and a continuous transition becomes a discontinuous one with hysteresis. The tipping point was a property of the assumed topology.
Then there is whether you could recover the parameters even with data. Chowell and colleagues’ 2023 primer on structural identifiability demonstrates that in models no more complex than SEIR, transmission rate and population size are correlated to the point of being unidentifiable from incidence data alone — multiple parameter combinations produce identical observations. Their warning is that “ignoring the structural identifiability of the model parameters can lead to erroneous inferences.” A 2021 Nature Human Behaviour review of models that fold in social and behavioural factors noted that adaptive network models tend to pick their behavioural terms by “an ‘Occam’s razor’ approach — incorporating intuitive mechanisms of which the mathematical form is often chosen for convenience.” That is the water this paper swims in, and it is not a criticism unique to these nine authors.
What we do have on the cognitive question is thin enough to describe in a paragraph.
The study everyone reaches for is Kosmyna and colleagues’ EEG work on essay writing with ChatGPT — weaker neural connectivity in the LLM arm, reduced sense of ownership, essays two teachers called soulless. It is worth knowing three things about it before leaning on it: 54 people enrolled and only 18 completed the crossover session, it was first posted in June 2025, and it has not been peer reviewed. A formal comment paper has since flagged its sample size, reproducibility and EEG methodology and suggested the results “could be interpreted more conservatively.”
Against that, a mostly Bocconi and OpenAI team — Asirvatham, Betti, Camuffo, Chatterji, Gambardella and seven co-authors — ran a 2×2 randomised trial with 1,053 first-year undergraduates in August 2026 and found the opposite sign on the outcome they measured: ChatGPT access improved evaluated scores, roughly half of it through greater coherence and more ideas. Neither study measures a transition rate. They are not even measuring the same construct. This is the same instrument problem I keep coming back to in measuring progress toward AGI — the number that gets quoted is rarely the number that was measured.
There is a version of this done properly, and it is old. Frank Bass published his diffusion model in 1969 with two parameters, p for external influence and q for word of mouth, and he fitted them to historical sales series for eleven consumer durables; later work extended the validation to hundreds of product categories. The parameters were estimated from data and the model was then tested against outcomes it had not seen. That is what earns a threshold.
It is not only a 1969 standard, either. Chisholm and colleagues built a competing-infection model of a divisive idea spreading through a population, with support and scepticism as rival contagions, and fitted its parameters to polling from the 2016 US Republican primary. Their conclusion is modest — that the model is “plausible for the spread of viewpoints of a divisive idea” — and it is modest because it was checked. Same family of model, with the parameters pinned to something that happened.
The cautionary tale is more recent and closer to the bone. The social-contagion literature of the late 2000s — obesity and happiness spreading through friendship networks — was met by Cohen-Cole and Fletcher and, separately, Russell Lyons, who showed that the same methods made implausible traits look contagious and that the effects were not distinguishable from shared environment once standard corrections were applied. The lesson was not that contagion is never real. It was that a contagion model fitted to observational data cannot, on its own, separate transmission from people simply resembling their friends.
I do not think this paper is doing anything wrong. It says what it is: a modelling exercise that identifies a structural possibility, plus the inverse condition it calls cognitive immunization. Papers like this are how you find out which questions are worth measuring.
The failure mode is downstream, and it is predictable. “Abrupt losses in cognitive competence past a critical threshold” is a sentence that will appear in policy briefs and op-eds within a month, and the four numbers that produced the threshold — 0.10, 0.40, 0.20, 0.10 — will not travel with it. A tipping point is what this class of model does when you give it a nonlinear reinforcement term. Getting one out is not evidence that one exists.
What would change my mind is a panel: the same several thousand people, followed for two years, with a pre-registered operational definition of each compartment and a measured count of how many crossed each boundary each quarter. That study would give you λ, ρ, μ and σ with confidence intervals, and then the model would be worth running. Until someone does it, the reading that holds is the one the authors already wrote down — an illustrative modelling assumption, not a general claim about LLM use.
References
Hacker News figures are a snapshot taken on 6 September 2026 and were still moving.