← Gautam Parab

The Cognitive Virus Model Has Five Rates. None of Them Has Been Measured.

The phrase that travelled was “abrupt losses in cognitive competence.” It comes from the abstract of “Large-Language Models as a Cognitive Virus,” posted to arXiv on 3 September 2026 by nine authors including David Krakauer, who runs the Santa Fe Institute, and Michael Levin at Tufts. Two days later it was on the front page of Hacker News, where it had collected 268 points and 196 comments when I looked. It is a good paper with a memorable title, and the memorable title is going to outrun the paper.

Here is what the model is. Three compartments: U, uncoupled users with little or no LLM use; C, coupled-autonomous users who use the tools regularly but retain independent cognition; and D, persistently dependent users for whom the interaction dominates. People move between them at rates — λ from U to C, ρ back again, μ from C down to D, σ from D back up to C — plus a fifth parameter, κ, weighting a nonlinear term that runs the other way. Each compartment carries a competence weight, and the paper sets those to 1.0, 0.5 and 0.1.

κ is worth pausing on, because it is the parameter that resists the headline. The term κU²C is added to the uncoupled population and subtracted from the coupled one, which makes it a second return flow alongside ρ, scaling with the square of how many uncoupled people there already are. The authors call it an Allee-like cooperative mechanism: independent cognitive practice is socially reinforced, so “cultural expectations that reward independent reasoning become more effective when autonomous individuals are common.” Raising κ makes adoption harder to establish, not easier.

The three-compartment model and its five rate parameters Three boxes in a row: uncoupled users, coupled autonomous users, and persistently dependent users. Four arrows connect them: transmission lambda from uncoupled to coupled, return rho from coupled back to uncoupled, progression mu from coupled to dependent, and recovery sigma from dependent back to coupled. A fifth parameter, kappa, weights a nonlinear Allee-like term, kappa times U squared times C, which acts as a second return flow from coupled back to uncoupled and therefore resists adoption rather than accelerating it. Each box carries an illustrative cognitive-competence weight: 1.0 for uncoupled, 0.5 for coupled, 0.1 for dependent. A footnote records that the paper's figures assign fixed values to rho, kappa, mu and sigma and fit none of them to data. MODEL STRUCTURE · SOLÉ ET AL., ARXIV:2609.03344, 3 SEP 2026 κU²C — second return flow, resists adoption Uncoupled Coupled Dependent little or no use regular, autonomous interaction dominates competence 1.0 0.5 0.1 λ ρ μ σ Five rate parameters. The published figures fix ρ = 0.10, κ = 0.40, μ = 0.20, σ = 0.10 and sweep λ. The values are labelled “parameters used.” There is no fitting procedure, no dataset, and no citation to a measured rate, because no measured rate exists. The competence weights 1.0 / 0.5 / 0.1 are described by the authors as illustrative. Compartment names and transitions as given in the paper; the layout is mine.
The structure is standard epidemiology. What epidemiology also has, and this does not, is a case count.

That is the whole apparatus, and the apparatus is fine. Compartmental models are a respectable way to reason about how something spreads, and this one is put together by people who know what they are doing — Krakauer at Santa Fe, Levin at Tufts, Ricard Solé at ICREA and Pompeu Fabra, Manlio De Domenico at Padua. The paper is also honest about its own status. When it orders the compartments by competence it says plainly: “This ordering is an illustrative modelling assumption, not a general claim about LLM use.” It concedes that its mean-field maths sits on top of a network topology that real populations do not have.

The problem is the distance between that sentence and the sentence about abrupt losses in cognitive competence, and which of the two gets quoted.

Nobody has measured a single rate

Run down the five parameters and ask what it would take to put a number on any of them.

σ, the recovery rate from dependence back to autonomy, is the one I keep getting stuck on. To measure it you need to define a recovery event: a person who was persistently dependent and now is not. Over what interval? Dependent on what operational test — self-report, task performance without the tool, some behavioural trace? Measured against what baseline, given that the same person’s unaided competence was never recorded before they started? Epidemiology gets σ from people who were sick and then were not, and it knows which is which because there is a clinical definition of both states. There is no clinical definition of cognitive dependence on a chatbot, which means there is no denominator, which means there is no rate. It is the same hole I wrote about in what per-mile crash statistics can and cannot show: the number is unavailable because the category has not been operationalised.

μ has the same problem in the other direction. κ is worse — it asks how strongly the prevalence of autonomous people reinforces someone else’s return to autonomy, which would require watching the same population at different compositions and measuring how the return rate changes. λ is the only one with a plausible path to measurement, and even there the available data measures the wrong thing.

Surveys measure stocks; the model needs flows Two panels. The left panel, headed what surveys measure, shows four measured stock quantities from the Pew Research Center survey of 5,119 US adults published 17 June 2026: 49 percent of US adults ever use AI chatbots, 44 percent have used ChatGPT specifically, 38 percent of employed adults use chatbots for work tasks, and 24 percent use chatbots daily. The right panel, headed what the model needs, lists the five rate parameters lambda, rho, mu, sigma and kappa, each shown as an empty dashed slot with no value, because no transition rate between levels of AI dependence has been measured in any population. STOCKS VS FLOWS · PEW, 17 JUN 2026, N=5,119 · MODEL PARAMETERS, SOLÉ ET AL. 2026 What surveys measure: stocks What the model needs: flows ever use a chatbot used ChatGPT use one for work use one daily 49% of US adults ever use AI chatbots — Pew, 17 Jun 2026 44% have used ChatGPT — Pew, 17 Jun 2026 38% of employed adults use chatbots for work tasks — Pew, 17 Jun 2026 24% use a chatbot daily — Pew, 17 Jun 2026 49% 44% 38% 24% λ uncoupled → coupled ρ coupled → uncoupled μ coupled → dependent σ dependent → coupled κ reinforcement weight no measurement no measurement no measurement no measurement no measurement A stock is how many people are in a state right now. A flow is the rate at which they move between states. Only the left panel exists.
Every adoption number in circulation is on the left. Every number the model runs on is on the right.

Pew put the stocks on the record in June: 49% of US adults say they have used an AI chatbot, 44% have used ChatGPT, 24% use one daily, 38% of employed adults use one for work, from a sample of 5,119 surveyed in February. Those are real measurements and they tell you nothing about transitions. A cross-section counts who is in each compartment on one day. The model runs on how fast people move between compartments, and nobody has followed a panel of users long enough to say.

It is not that the states are unmeasurable. Li and colleagues published a validated Large Language Model Dependence Scale in 2025, an eighteen-item bifactor instrument separating what they call functional dependence from existential dependence, developed on an exploratory sample of 421 users and a confirmatory sample of 1,030. That is a serviceable definition of compartment D. Point it at the same people every quarter for two years and you would have μ and σ. Nobody has, so the instrument measures a state and the model still runs on flows.

The tipping point is a property of the form

The deeper problem is that this class of model produces thresholds whether or not the world has any.

Bissell and colleagues showed the mechanism cleanly in 2014, in a compartmental model of smoking. They took two rates that had previously been constant — the rate at which smokers quit, and the rate at which former smokers relapse — and made them depend on peer influence. Their own summary of the result is the whole argument in one sentence: it “not only modifies the number of steady-states and nature of their asymptotic stability, but also introduces a new kind of non-linear ‘tipping-point’ dynamic.” The tipping point arrived with the functional form, not with any new data.

Iacopini and colleagues made the same point from a different angle in Nature Communications in 2019: hold the parameters identical and switch the interactions from pairwise to higher-order, and a continuous transition becomes a discontinuous one with hysteresis. The tipping point was a property of the assumed topology.

Then there is whether you could recover the parameters even with data. Chowell and colleagues’ 2023 primer on structural identifiability demonstrates that in models no more complex than SEIR, transmission rate and population size are correlated to the point of being unidentifiable from incidence data alone — multiple parameter combinations produce identical observations. Their warning is that “ignoring the structural identifiability of the model parameters can lead to erroneous inferences.” A 2021 Nature Human Behaviour review of models that fold in social and behavioural factors noted that adaptive network models tend to pick their behavioural terms by “an ‘Occam’s razor’ approach — incorporating intuitive mechanisms of which the mathematical form is often chosen for convenience.” That is the water this paper swims in, and it is not a criticism unique to these nine authors.

The empirical base is two studies pointing different ways

What we do have on the cognitive question is thin enough to describe in a paragraph.

The two studies most often invoked, side by side Two study panels. The Kosmyna and colleagues EEG essay-writing study, an arXiv preprint first posted 10 June 2025 and not peer reviewed, enrolled 54 participants of whom only 18 completed the fourth crossover session, and found weaker neural connectivity and reduced essay ownership among ChatGPT users. The Asirvatham and colleagues randomised controlled trial, dated August 2026, enrolled 1,053 undergraduates in a two-by-two design and found that ChatGPT access improved evaluated scores, with roughly half the effect attributable to greater coherence and idea count. A footnote notes the sample sizes differ by a factor of nearly twenty and that the two studies measure different things. THE MEASUREMENT BASE · TWO STUDIES, DIFFERENT DIRECTIONS Kosmyna et al., EEG essay writing arXiv preprint, first posted 10 Jun 2025 · not peer reviewed Asirvatham et al., randomised controlled trial working paper, Aug 2026 · 2×2 design, undergraduates 54 participants enrolled; 18 completed session 4 — Kosmyna et al., 2025 1,053 participants — Asirvatham et al., Aug 2026 54 enrolled, 18 finished the crossover 1,053 weaker neural connectivity, reduced sense of ownership access improved evaluated scores; about half via coherence and idea count The samples differ by nearly twenty times, and the studies do not measure the same construct: neural correlates during one writing task is not the same outcome as evaluated reasoning quality. Neither measures a transition rate.
The small, older, un-peer-reviewed study is the one that travels. The larger, newer randomised trial points the other way, and neither gives the model a parameter.

The study everyone reaches for is Kosmyna and colleagues’ EEG work on essay writing with ChatGPT — weaker neural connectivity in the LLM arm, reduced sense of ownership, essays two teachers called soulless. It is worth knowing three things about it before leaning on it: 54 people enrolled and only 18 completed the crossover session, it was first posted in June 2025, and it has not been peer reviewed. A formal comment paper has since flagged its sample size, reproducibility and EEG methodology and suggested the results “could be interpreted more conservatively.”

Against that, a mostly Bocconi and OpenAI team — Asirvatham, Betti, Camuffo, Chatterji, Gambardella and seven co-authors — ran a 2×2 randomised trial with 1,053 first-year undergraduates in August 2026 and found the opposite sign on the outcome they measured: ChatGPT access improved evaluated scores, roughly half of it through greater coherence and more ideas. Neither study measures a transition rate. They are not even measuring the same construct. This is the same instrument problem I keep coming back to in measuring progress toward AGI — the number that gets quoted is rarely the number that was measured.

Bass had eleven product histories

There is a version of this done properly, and it is old. Frank Bass published his diffusion model in 1969 with two parameters, p for external influence and q for word of mouth, and he fitted them to historical sales series for eleven consumer durables; later work extended the validation to hundreds of product categories. The parameters were estimated from data and the model was then tested against outcomes it had not seen. That is what earns a threshold.

It is not only a 1969 standard, either. Chisholm and colleagues built a competing-infection model of a divisive idea spreading through a population, with support and scepticism as rival contagions, and fitted its parameters to polling from the 2016 US Republican primary. Their conclusion is modest — that the model is “plausible for the spread of viewpoints of a divisive idea” — and it is modest because it was checked. Same family of model, with the parameters pinned to something that happened.

The cautionary tale is more recent and closer to the bone. The social-contagion literature of the late 2000s — obesity and happiness spreading through friendship networks — was met by Cohen-Cole and Fletcher and, separately, Russell Lyons, who showed that the same methods made implausible traits look contagious and that the effects were not distinguishable from shared environment once standard corrections were applied. The lesson was not that contagion is never real. It was that a contagion model fitted to observational data cannot, on its own, separate transmission from people simply resembling their friends.

What I would want before believing the threshold

I do not think this paper is doing anything wrong. It says what it is: a modelling exercise that identifies a structural possibility, plus the inverse condition it calls cognitive immunization. Papers like this are how you find out which questions are worth measuring.

The failure mode is downstream, and it is predictable. “Abrupt losses in cognitive competence past a critical threshold” is a sentence that will appear in policy briefs and op-eds within a month, and the four numbers that produced the threshold — 0.10, 0.40, 0.20, 0.10 — will not travel with it. A tipping point is what this class of model does when you give it a nonlinear reinforcement term. Getting one out is not evidence that one exists.

What would change my mind is a panel: the same several thousand people, followed for two years, with a pre-registered operational definition of each compartment and a measured count of how many crossed each boundary each quarter. That study would give you λ, ρ, μ and σ with confidence intervals, and then the model would be worth running. Until someone does it, the reading that holds is the one the authors already wrote down — an illustrative modelling assumption, not a general claim about LLM use.


References

  1. Solé, Ruffini, Castaldo, Tuccio, Seoane, De Domenico, Elena, Krakauer, & Levin. (2026, September 3). “Large-Language Models as a Cognitive Virus”. arXiv:2609.03344.
  2. Pew Research Center. (2026, June 17). “Americans and AI 2026: Chatbots, Smart Devices and Views on Impact”. N=5,119, fielded 17–23 February 2026.
  3. Kosmyna et al. (2025, June 10, v1). arXiv:2506.08872. Preprint, not peer reviewed.
  4. Comment paper responding to Kosmyna et al. arXiv:2601.00856.
  5. Asirvatham, Betti, Brown, Camuffo, Chatterji, Fumagalli, Gambardella, Mariani, Pandey, Ramos, Salvucci, & Simic. (2026, August). “Training Novices to Think, or Giving Them LLMs? Evidence from an RCT”. CEPR DP21882 / OpenAI working paper. N=1,053.
  6. Li et al. (2025). “From assistance to reliance: development and validation of the large language model dependence scale.” International Journal of Information Management. DOI: 10.1016/j.ijinfomgt.2025.102888.
  7. Bissell et al. (2014). Mathematical Models and Methods in Applied Sciences. DOI: 10.1142/s0218202513500656. Methodological/historical context, pre-dating the nine-month research window.
  8. Iacopini et al. (2019). Nature Communications. DOI: 10.1038/s41467-019-10431-6. Methodological/historical context, pre-dating the nine-month research window.
  9. Chowell et al. (2023). Journal of Mathematical Biology. DOI: 10.1007/s00285-023-02007-2. Methodological/historical context, pre-dating the nine-month research window.
  10. Bedson et al. (2021). Nature Human Behaviour. DOI: 10.1038/s41562-021-01136-2. Methodological/historical context, pre-dating the nine-month research window.
  11. Chisholm et al. (2019, published online 2018). Journal of Mathematical Sociology. DOI: 10.1080/0022250x.2018.1555828. Methodological/historical context, pre-dating the nine-month research window.
  12. Bass, F. “A New Product Growth for Model Consumer Durables.” Management Science (1969). Historical.
  13. Cohen-Cole & Fletcher. (2008). On social-contagion identification. Historical.
  14. Lyons. (2011). On social-contagion identification. Historical.

Hacker News figures are a snapshot taken on 6 September 2026 and were still moving.