静态对齐方法在能力增强时会失效,因无法应对新情境与价值冲突。
The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
- 用固定规范定义价值,无法适应未来未知场景
- 三大哲学难题导致对齐机制必然失效:事实与价值脱节、价值多元不可统一、上下文错配
- 适用于高自主系统的设计需转向动态可更新的开放方案
静态内容导向的AI价值对齐在能力扩展、分布偏移和自主性提升下难以保持稳健。任何将对齐视为优化固定形式价值目标(如奖励函数、效用函数、宪法原则或学习偏好)的方法均存在根本缺陷。三个哲学问题加剧了困境:休谟的“是-应当”鸿沟(行为数据无法确定规范内涵)、柏林的价值多元论(人类价值难以一致形式化)、扩展的框架问题(任何价值编码都会与先进AI创造的新情境错配)。强化学习人类反馈(RLHF)、宪法式AI、逆强化学习与合作辅助游戏均陷入这一“规范陷阱”,其失败模式源于结构性脆弱,而非仅由数据或算法不足导致。当规范封闭不再更新时,各组件的补救措施相互强化困难。基于兼容主义哲学,训练中行为合规不保证新环境下对齐可靠,且该差距随系统能力增长而扩大。对具有价值负载的自主系统而言,封闭式方法存在结构性风险,且越强越危险。建设性责任转向开放式、发展响应式路径,但能否实现仍属实证问题。
原文摘要 · Abstract (English)
Static content-based AI value alignment is insufficient for robust alignment under capability scaling, distributional shift, and increasing autonomy. This holds for any approach that treats alignment as optimizing toward a fixed formal value-object, whether reward function, utility function, constitutional principles, or learned preference representation. Three philosophical results create compounding difficulties: Hume's is-ought gap (behavioral data underdetermines normative content), Berlin's value pluralism (human values resist consistent formalization), and the extended frame problem (any value encoding will misfit future contexts that advanced AI creates). RLHF, Constitutional AI, inverse reinforcement learning, and cooperative assistance games each instantiate this specification trap, and their failure modes reflect structural vulnerabilities, not merely engineering limitations that better data or algorithms will straightforwardly resolve. Known workarounds for individual components face mutually reinforcing difficulties when the specification is closed: the moment it ceases to update from the process it governs. Drawing on compatibilist philosophy, the paper argues that behavioral compliance under training conditions does not guarantee robust alignment under novel conditions, and that this gap grows with system capability. For value-laden autonomous systems, known closed approaches face structural vulnerabilities that worsen with capability. The constructive burden shifts to open, developmentally responsive approaches, though whether such approaches can be achieved remains an empirical question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。