Paper 2601.19062

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

Published
Jan 2026
Research lab
Anthropic
Citations
19
GitHub
Not linked

01 In brief

Summary

This paper presents the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude.ai conversations using a privacy-preserving approach.

The authors define situational disempowerment as occurring when interactions risk leading users to distorted perceptions of reality, inauthentic value judgments, or actions misaligned with their values.

They measure three disempowerment potential primitives (reality, value judgment, and action distortion) and four amplifying factors (authority projection, attachment, reliance/dependency, vulnerability).

Quantitatively, severe disempowerment potential occurs in fewer than one in a thousand conversations, but rates are higher in personal domains like relationships and lifestyle.

Qualitatively, they uncover patterns such as validation of persecution narratives and grandiose identities with sycophantic language, definitive moral judgments about third parties, and complete scripting of value-laden communications.

Historical analysis shows an increase in disempowerment potential over time, and interactions with greater disempowerment potential receive higher user approval ratings.

Preference models trained on human feedback do not robustly disincentivize disempowerment.

The findings highlight a tension between short-term user preferences and long-term human empowerment, motivating AI systems designed to support human autonomy.

02 From the paper

Abstract

Although AI assistants are now deeply embedded in society, there has been limited empirical study of how their usage affects human empowerment. We present the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude$.$ai conversations using a privacy-preserving approach. We focus on situational disempowerment potential, which occurs when AI assistant interactions risk leading users to form distorted perceptions of reality, make inauthentic value judgments, or act in ways misaligned with their values. Quantitatively, we find that severe forms of disempowerment potential occur in fewer than one in a thousand conversations, though rates are substantially higher in personal domains like relationships and lifestyle. Qualitatively, we uncover several concerning patterns, such as validation of persecution narratives and grandiose identities with emphatic sycophantic language, definitive moral judgments about third parties, and complete scripting of value-laden personal communications that users appear to implement verbatim. Analysis of historical trends reveals an increase in the prevalence of disempowerment potential over time. We also find that interactions with greater disempowerment potential receive higher user approval ratings, possibly suggesting a tension between short-term user preferences and long-term human empowerment. Our findings highlight the need for AI systems designed to robustly support human autonomy and flourishing.