Google DeepMind published version 3.1 of its Frontier Safety Framework on April 17, 2026. It adds a risk domain the earlier versions did not track, Harmful Manipulation, and defines the capability threshold this way:
Possesses manipulative capabilities sufficient to enable it to systematically and substantially change beliefs and behavior in identified high stakes contexts over the course of interactions with the model, reasonably resulting in additional expected harm at severe scale.
Read it again with an eye on one word. Beliefs and behavior. The conjunction carries the whole assumption: it treats the two as one capability, sized by one threshold, crossed at one moment. If a model can move what you think, it can move what you do, and a single assessment can tell us when that has happened.
Seven days earlier, a preprint measured both in the same experiments and found they do not correlate.
What happens when you measure both
The paper is Artificial intelligence can persuade people to take political actions, submitted to arXiv on April 10,
- It runs two preregistered experiments, gathering 17,950 responses from 14,779 UK adults: 8,000 people in the first study and 9,950 in the second. In both, conversational models try to shift participants on eight political causes.
The design’s whole point is that it does not stop at the survey. Alongside attitude scales, it records whether participants signed a real petition and whether they gave money. The behavioral effects are large. In Study 1, treatment participants were 12.8 percentage points more likely to sign; in Study 2, 19.7 points, both at p < .001. If you wanted a headline confirming that AI persuasion reaches past opinion and into action, that is the number.
Then the authors plotted one against the other. Each point is one experimental condition. Study 1 has fifteen of them, five models crossed with three conversation types; Study 2 has the eight persuasion strategies. The x-axis is how much that condition moved support for the petition. The y-axis is how much it moved signing the petition.
In Study 1 the correlation is r = 0.05, p = 0.85. In Study 2, across the eight strategies, r = 0.18, p = 0.68. Neither is distinguishable from nothing. A condition’s power to move opinion tells you essentially nothing about its power to move a hand toward a signature.
This is not a null result from an underpowered study. The same experiments produced significant, sizable effects on both outcomes separately. What failed to appear is the relationship between them.
Two outcomes, two mechanisms
Two conditions could be uncorrelated by noise. The paper goes further and shows the two outcomes respond to different things.
Among the strategies tested, the one built on providing issue information produces the largest attitude effects, and no comparable advantage on behavior.
The same split shows up in what participants said they got out of the conversation. The paper asked them, on a 0–100 scale, how much they agreed that they had learned something new. Self-reported learning tracks the attitude effects almost perfectly, and tracks the behavioral effects not at all.
The authors put the implication plainly in their own abstract: previous findings relying on attitudinal outcomes “may generalize poorly to behaviour, and therefore risk substantially mischaracterizing the real-world behavioural impact of AI persuasion.” It is an unusual thing for a paper to do: report a strong positive headline and, in the same breath, tell its own field that its dominant instrument may have been measuring the wrong thing.
The instrument has been the attitude scale
That dominance is well documented. Claire Wardle, Pete Brown and David Scales, writing in Media and Communication on July 30, 2026, describe the field as having “relied predominantly on experimental methods rooted in social psychology, operationalizing persuasion as short-term attitudinal change measured through randomized controlled trials.” They argue that approach cannot explain how this kind of persuasion actually works.
Several 2026 studies do measure behavior. AI systems out-persuade expert humans, from June, ran 18,978 conversations with 6,923 people and ended with a field study in which AI was “nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children.” A separate April study of donation appeals found LLM-authored text drew significantly more money than human-authored text, while noting the absolute differences were small. These are real behaviors with real costs.
But notice what they are: a handful of studies, clustered in one year, several from the same group, and every one of them a snapshot. None reports a follow-up measurement. The only 2026 durability data I can point to sits on the other side of the divide: a within-subject study of fifty-three people whose moral judgments shifted after brief chatbot conversations, with the effect measured again two weeks later.
Fifty-three people is a small study and I would not build policy on it. What it illustrates is the shape of the evidence base: the durable measurement is self-report, the costly measurement is a snapshot, and nothing joins them. The same paper notes the shift “did not extend to punishment”: participants revised what they judged, not what they were willing to impose.
What the conjunction assumes
None of this makes the Frontier Safety Framework wrong to name harmful manipulation as a risk domain. I think naming it was right, and the document is more candid than most: it states that research here “is nascent,” that the threshold “is exploratory and subject to further research, and may be substantially changed over time.” That is an honest hedge, and a week between a preprint and a published framework is not a reading window. Nobody missed anything.
The problem is structural, and it survives the hedge. A threshold written as “beliefs and behavior” has to be assessed somehow, and whoever assesses it will reach for an instrument. If they reach for attitude scales, which are the field’s default and by far the best-supplied with prior art, they will produce a number that, on the best evidence available, carries no information about the behavioral half of the sentence. Measured in the same people, in the same conversations, and in the same week, the two came apart.
This is the same failure mode I keep running into from different directions. When California’s executive order asked whether a shutdown mechanism’s efficacy could be verified on an ongoing basis, the gap was that no measurement protocol existed for the thing being required. When a vendor advertises a speed multiple, the question is always what sits in the denominator. Here the gap is narrower and stranger: the measurement protocols exist, several of them, and they disagree about what they are measuring.
The cheap fix is available and I would take it. Split the conjunction. Write two thresholds, name the instrument for each, and require that a model assessed against the behavioral one be assessed with a behavioral outcome: something a participant pays for in money, time, or signature. It costs more. That is the entire objection to it, and it is not a good one, because the alternative is a threshold that can be cleared or missed on a survey score which, in the experiments that measured both, predicted nothing about the behaviour the domain exists to prevent.
Until then, the “and” is not describing a capability. It is describing two, and hoping they travel together.
References
- Google DeepMind (2026). Frontier Safety Framework, Version 3.1. Published 17 April 2026.
- Hackenburg et al. (2026). Artificial intelligence can persuade people to take political actions. arXiv:2604.09200, 10 April 2026.
- Hackenburg et al. (2026). AI systems out-persuade expert humans. arXiv:2606.16475, 15 June 2026.
- Wardle, C., Brown, P., & Scales, D. (2026). A Research Agenda for Studying LLM-Powered Persuasion Technology to Mitigate Misinformation. Media and Communication 14, Article 12442, 30 July 2026. DOI 10.17645/mac.12442.
- Caffier, J., Stavrova, O., & Kleinberg, B. (2026). Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization. arXiv:2604.03202, 3 April 2026 (revised 18 July 2026).
- Teng, Y., Zhong, Q., Nguyen Thordsen, K. M. T., Montag, C., & Becker, B. (2026). Brief chatbot interactions produce lasting changes in human moral values. arXiv:2604.21430, 23 April 2026.