AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Gölz et al. (2025) demonstrated that the \textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $β$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distor

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:22:04.777Z. This is not the publication date.