arXiv:2608.24046cs.AI2026-08

用福利影响重写对齐问题,让AI决策更公平可设计。

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

论文配图:Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
图 1 · 摘自论文原文
  • 将对齐转化为福利影响的线性优化,引入经济学工具
  • 证明投票和随机独裁机制防策略操纵且一致
  • 可定制公平约束,适合关注伦理与公平的开发者

当AI决策影响多人时,对齐本质上是社会选择问题:如何调和不同人的偏好并形成统一模型?当前主流的基于人类反馈强化学习方法忽视了这一核心问题,缺乏社会选择保障。本文提出新范式:聚焦算法的福利影响,将对齐重构为凸影响空间上的线性优化,使福利经济学与机制设计工具得以应用。该框架揭示对齐协议如何转化为福利后果,并反向指导协议设计。研究证明,按议题投票与随机独裁机制具有策略不变性和一致性。进一步,基于影响表示推导出一系列对齐协议,在满足个体或群体伤害上限等社会目标下最大化功利主义福利。通过真实人类偏好数据(肾脏分配、慈善食物分发、LLM回应、电车难题)验证了不同对齐方案的福利影响。

原文摘要 · Abstract (English)

When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\unicode{x2013}$reinforcement learning from human feedback$\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.

AI对齐社会选择福利优化公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。