arXiv:2607.00001cs.AIcs.CY2026-07AAAI被引 2

AI对人类偏好的影响是动态的,该研究提出新范式控制这种演变过程。

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

论文配图:Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
图 1 · 摘自论文原文
  • 将偏好视为互动中不断演化的过程,而非固定目标
  • 用控制理论建模人与AI交互如何改变价值判断
  • 适合关注长期人机价值共塑的研究者

现有AI对齐方法将人类偏好视为需推断和优化的静态目标,但大量实证证据表明偏好是分层、动态且通过互动构建的,尤其在面对自适应技术时更为明显。随着AI系统日益持久、个性化和嵌入社会,它们正越来越多地参与塑造人们关注、重视和认可的内容。本文提出‘建构性对齐’(Constructive Alignment)新范式,将对齐问题重新定义为对人类偏好轨迹演化的控制,而非静态偏好满足。基于行为经济学、心理学与建构主义社会理论,我们把偏好建模为分层的状态变量,在与AI系统的交互中持续演化。通过控制论框架,系统行为与交互设计共同影响世界状态与人类评价状态。我们认为,对齐的核心并非控制AI行为,而是调控其如何影响偏好演化——确保价值轨迹具有连贯性、反思性认同、认识论基础,避免操纵,并在不确定性下增强自主性。因此,对齐本质上成为治理长期价值形成的问题。

原文摘要 · Abstract (English)

Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.

AI对齐偏好建模人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。