Google's Manipulation Threshold Says 'Beliefs and Behavior.' The Measurements Don't Correlate.

Google DeepMind published version 3.1 of its Frontier Safety Framework on April 17, 2026. It adds a risk domain the earlier versions did not track, Harmful Manipulation, and defines the capability threshold this way:

Possesses manipulative capabilities sufficient to enable it to systematically and substantially change beliefs and behavior in identified high stakes contexts over the course of interactions with the model, reasonably resulting in additional expected harm at severe scale.

Read it again with an eye on one word. Beliefs and behavior. The conjunction carries the whole assumption: it treats the two as one capability, sized by one threshold, crossed at one moment. If a model can move what you think, it can move what you do, and a single assessment can tell us when that has happened.

Seven days earlier, a preprint measured both in the same experiments and found they do not correlate.

What happens when you measure both

The paper is Artificial intelligence can persuade people to take political actions, submitted to arXiv on April 10,

  1. It runs two preregistered experiments, gathering 17,950 responses from 14,779 UK adults: 8,000 people in the first study and 9,950 in the second. In both, conversational models try to shift participants on eight political causes.

The design’s whole point is that it does not stop at the survey. Alongside attitude scales, it records whether participants signed a real petition and whether they gave money. The behavioral effects are large. In Study 1, treatment participants were 12.8 percentage points more likely to sign; in Study 2, 19.7 points, both at p < .001. If you wanted a headline confirming that AI persuasion reaches past opinion and into action, that is the number.

Then the authors plotted one against the other. Each point is one experimental condition. Study 1 has fifteen of them, five models crossed with three conversation types; Study 2 has the eight persuasion strategies. The x-axis is how much that condition moved support for the petition. The y-axis is how much it moved signing the petition.

Two scatter plots stacked vertically, black on white. The top panel, labelled Study 1 with 15 conditions, plots treatment effect on supporting a petition against treatment effect on signing it; the points scatter without pattern and the dashed fit line is nearly flat, annotated r = 0.05, p = 0.85. The lower panel, Study 2 with 8 strategies, shows the same flat scatter annotated r = 0.18, p = 0.68. Both fits carry wide grey confidence bands that widen at the edges.
The conditions that changed minds most are not the conditions that got signatures. Each dot is one experimental condition; the dashed line is the linear fit and the grey band its 95 percent confidence interval. Neither slope is distinguishable from flat. Image: Hackenburg et al., "Artificial intelligence can persuade people to take political actions", arXiv:2604.09200v1, Figure 3A, CC BY 4.0. Cropped to panel A.

In Study 1 the correlation is r = 0.05, p = 0.85. In Study 2, across the eight strategies, r = 0.18, p = 0.68. Neither is distinguishable from nothing. A condition’s power to move opinion tells you essentially nothing about its power to move a hand toward a signature.

This is not a null result from an underpowered study. The same experiments produced significant, sizable effects on both outcomes separately. What failed to appear is the relationship between them.

Two outcomes, two mechanisms

Two conditions could be uncorrelated by noise. The paper goes further and shows the two outcomes respond to different things.

Among the strategies tested, the one built on providing issue information produces the largest attitude effects, and no comparable advantage on behavior.

Information prompts raise attitude effects but not behavioural effects Eight paired comparisons from Hackenburg et al. 2026, each showing the average treatment effect for a non-information prompt and for an information prompt. Among attitude outcomes, measured in standardized units, every pair rises with information: Study 1 support petition 0.07 to 0.20, Study 1 support organization 0.13 to 0.25, Study 2 support petition 0.13 to 0.25, Study 2 support organization 0.16 to 0.32. Among behavioural outcomes, measured as proportions, the pairs barely move and two fall: Study 1 sign petition 0.13 to 0.12, Study 1 organizational engagement 0.06 to 0.07, Study 2 sign petition 0.20 to 0.16, Study 2 organizational engagement 0.06 to 0.09. AVERAGE TREATMENT EFFECT · INFORMATION PROMPT VS OTHER · HACKENBURG ET AL., APR 2026 ATTITUDE OUTCOMES · STANDARDIZED CHANGE SCORES BEHAVIOURAL OUTCOMES · PROPORTION SIGNING OR ENGAGING Study 1 · Support petition Study 1 · Support organization Study 2 · Support petition Study 2 · Support organization Study 1 · Sign petition Study 1 · Organizational engagement Study 2 · Sign petition Study 2 · Organizational engagement 0.07 — other prompts, Study 1 support petition 0.13 — other prompts, Study 1 support organization 0.13 — other prompts, Study 2 support petition 0.16 — other prompts, Study 2 support organization 0.13 — other prompts, Study 1 sign petition 0.06 — other prompts, Study 1 organizational engagement 0.20 — other prompts, Study 2 sign petition 0.06 — other prompts, Study 2 organizational engagement 0.20 — information prompt, Study 1 support petition 0.25 — information prompt, Study 1 support organization 0.25 — information prompt, Study 2 support petition 0.32 — information prompt, Study 2 support organization 0.12 — information prompt, Study 1 sign petition 0.07 — information prompt, Study 1 organizational engagement 0.16 — information prompt, Study 2 sign petition 0.09 — information prompt, Study 2 organizational engagement 0.07 → 0.20 0.13 → 0.25 0.13 → 0.25 0.16 → 0.32 0.13 → 0.12 0.06 → 0.07 0.20 → 0.16 0.06 → 0.09 0.0 0.1 0.2 0.3 Other prompts Information (Issue) prompt Attitude effects are standardized change scores, (post − pre) / SD(pre); behavioural effects are proportions. The two families share an axis but not a unit.
Information moves the top block and leaves the bottom block alone. Two of the four behavioural pairs move the wrong way.

The same split shows up in what participants said they got out of the conversation. The paper asked them, on a 0–100 scale, how much they agreed that they had learned something new. Self-reported learning tracks the attitude effects almost perfectly, and tracks the behavioral effects not at all.

Self-reported learning predicts attitude change but not behaviour Eight correlation coefficients between self-reported learning and treatment effect, from Hackenburg et al. 2026. For attitude outcomes all four are high and significant: Study 1 support petition r = 0.87, Study 1 support organization r = 0.93, Study 2 support petition r = 0.85, Study 2 support organization r = 0.89, each p below .01. For behavioural outcomes none reaches significance: Study 1 sign petition r = minus 0.15 at p = 0.59, Study 1 organizational engagement r = 0.44 at p = 0.10, Study 2 sign petition r = minus 0.16 at p = 0.70, Study 2 organizational engagement r = 0.63 at p = 0.10. CORRELATION OF SELF-REPORTED LEARNING WITH TREATMENT EFFECT · HACKENBURG ET AL., APR 2026 ATTITUDE OUTCOMES BEHAVIOURAL OUTCOMES Study 1 · Support petition Study 1 · Support organization Study 2 · Support petition Study 2 · Support organization Study 1 · Sign petition Study 1 · Organizational engagement Study 2 · Sign petition Study 2 · Organizational engagement r = 0.87, p < .01 — Study 1 support petition r = 0.93, p < .01 — Study 1 support organization r = 0.85, p < .01 — Study 2 support petition r = 0.89, p < .01 — Study 2 support organization r = -0.15, p = 0.59 — Study 1 sign petition r = 0.44, p = 0.10 — Study 1 organizational engagement r = -0.16, p = 0.70 — Study 2 sign petition r = 0.63, p = 0.10 — Study 2 organizational engagement 0.87 · p<.01 0.93 · p<.01 0.85 · p<.01 0.89 · p<.01 −0.15 · p=.59 0.44 · p=.10 −0.16 · p=.70 0.63 · p=.10 −0.3 0 0.3 0.7 1.0 Filled points are significant at p < .01; open points are not significant. Each point is one study-outcome pair.
Feeling informed is an excellent predictor of saying you agree, and no predictor at all of doing anything.

The authors put the implication plainly in their own abstract: previous findings relying on attitudinal outcomes “may generalize poorly to behaviour, and therefore risk substantially mischaracterizing the real-world behavioural impact of AI persuasion.” It is an unusual thing for a paper to do: report a strong positive headline and, in the same breath, tell its own field that its dominant instrument may have been measuring the wrong thing.

The instrument has been the attitude scale

That dominance is well documented. Claire Wardle, Pete Brown and David Scales, writing in Media and Communication on July 30, 2026, describe the field as having “relied predominantly on experimental methods rooted in social psychology, operationalizing persuasion as short-term attitudinal change measured through randomized controlled trials.” They argue that approach cannot explain how this kind of persuasion actually works.

Several 2026 studies do measure behavior. AI systems out-persuade expert humans, from June, ran 18,978 conversations with 6,923 people and ended with a field study in which AI was “nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children.” A separate April study of donation appeals found LLM-authored text drew significantly more money than human-authored text, while noting the absolute differences were small. These are real behaviors with real costs.

But notice what they are: a handful of studies, clustered in one year, several from the same group, and every one of them a snapshot. None reports a follow-up measurement. The only 2026 durability data I can point to sits on the other side of the divide: a within-subject study of fifty-three people whose moral judgments shifted after brief chatbot conversations, with the effect measured again two weeks later.

The only durability data is on self-report, and the effect grew A slope chart of Cohen's d from a 53-person study of chatbot influence on moral judgements. Immediately after the conversation the reported effect sizes span 0.735 to 1.576. At a two-week follow-up the span has risen to 1.038 to 2.069. Both the lower and upper bounds increase over the interval. COHEN'S D, MORAL-JUDGEMENT SHIFT · N = 53 · ARXIV:2604.21430, APR 2026 0.5 1.0 1.5 2.0 d = 0.735 — smallest immediate effect d = 1.038 — smallest effect at two weeks d = 1.576 — largest immediate effect d = 2.069 — largest effect at two weeks 0.735 1.576 1.038 2.069 Immediately after Two weeks later Self-reported moral judgements only. The effects did not extend to punishment decisions, and none of the behavioural studies above reports any follow-up.
The measured shift did not decay over two weeks; it widened. Whether anything people did followed the same curve is not a question the literature has answered. The paper's abstract and its results section give different upper bounds, 2.069 and 2.269; the figure uses the abstract's.

Fifty-three people is a small study and I would not build policy on it. What it illustrates is the shape of the evidence base: the durable measurement is self-report, the costly measurement is a snapshot, and nothing joins them. The same paper notes the shift “did not extend to punishment”: participants revised what they judged, not what they were willing to impose.

What the conjunction assumes

None of this makes the Frontier Safety Framework wrong to name harmful manipulation as a risk domain. I think naming it was right, and the document is more candid than most: it states that research here “is nascent,” that the threshold “is exploratory and subject to further research, and may be substantially changed over time.” That is an honest hedge, and a week between a preprint and a published framework is not a reading window. Nobody missed anything.

The problem is structural, and it survives the hedge. A threshold written as “beliefs and behavior” has to be assessed somehow, and whoever assesses it will reach for an instrument. If they reach for attitude scales, which are the field’s default and by far the best-supplied with prior art, they will produce a number that, on the best evidence available, carries no information about the behavioral half of the sentence. Measured in the same people, in the same conversations, and in the same week, the two came apart.

This is the same failure mode I keep running into from different directions. When California’s executive order asked whether a shutdown mechanism’s efficacy could be verified on an ongoing basis, the gap was that no measurement protocol existed for the thing being required. When a vendor advertises a speed multiple, the question is always what sits in the denominator. Here the gap is narrower and stranger: the measurement protocols exist, several of them, and they disagree about what they are measuring.

The cheap fix is available and I would take it. Split the conjunction. Write two thresholds, name the instrument for each, and require that a model assessed against the behavioral one be assessed with a behavioral outcome: something a participant pays for in money, time, or signature. It costs more. That is the entire objection to it, and it is not a good one, because the alternative is a threshold that can be cleared or missed on a survey score which, in the experiments that measured both, predicted nothing about the behaviour the domain exists to prevent.

Until then, the “and” is not describing a capability. It is describing two, and hoping they travel together.

References

  1. Google DeepMind (2026). Frontier Safety Framework, Version 3.1. Published 17 April 2026.
  2. Hackenburg et al. (2026). Artificial intelligence can persuade people to take political actions. arXiv:2604.09200, 10 April 2026.
  3. Hackenburg et al. (2026). AI systems out-persuade expert humans. arXiv:2606.16475, 15 June 2026.
  4. Wardle, C., Brown, P., & Scales, D. (2026). A Research Agenda for Studying LLM-Powered Persuasion Technology to Mitigate Misinformation. Media and Communication 14, Article 12442, 30 July 2026. DOI 10.17645/mac.12442.
  5. Caffier, J., Stavrova, O., & Kleinberg, B. (2026). Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization. arXiv:2604.03202, 3 April 2026 (revised 18 July 2026).
  6. Teng, Y., Zhong, Q., Nguyen Thordsen, K. M. T., Montag, C., & Becker, B. (2026). Brief chatbot interactions produce lasting changes in human moral values. arXiv:2604.21430, 23 April 2026.