负向约束比正向偏好更利于AI对齐,因其结构更稳定可靠。
Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences
- 用否定信号训练模型,避免正向偏好带来的表面迎合问题。
- 负向约束可形成离散、独立验证的明确边界,收敛更稳定。
- 适合关注安全对齐、避免模型讨好人类的研究者参考。
近期实证研究表明,仅使用负向反馈训练大语言模型可达到甚至超过标准的人类反馈强化学习(RLHF)效果。负样本强化学习在数学推理任务上与PPO方法持平;分布排斥优化仅用被排斥样本即可有效训练;宪法AI在安全性基准上优于纯RLHF。然而,尚无统一理论解释负向信号为何高效。本文提出:正向偏好与负向约束在结构上存在不对称性。正向偏好(‘哪个更好’)表达连续、依赖上下文的人类价值,无法穷尽描述,导致模型学习表面相关特征(如迎合用户);而负向约束(‘哪里错了’)是离散、有限、可独立验证的禁止项,能收敛至稳定边界。这一不对称性源于波普尔的可证伪逻辑与否定知识的认识论,解释了偏好强化中出现的讨好行为,并揭示了负向信号方法的有效性。我们主张,对齐研究应从‘学习人类偏好’转向‘学习人类拒绝’,并提出可检验的预测。
原文摘要 · Abstract (English)
Recent empirical results have demonstrated that training large language models (LLMs) with negative-only feedback can match or exceed standard reinforcement learning from human feedback (RLHF). Negative Sample Reinforcement achieves parity with PPO on mathematical reasoning; Distributional Dispreference Optimization trains effectively using only dispreferred samples; and Constitutional AI outperforms pure RLHF on harmlessness benchmarks. Yet no unified theoretical account explains why negative signals are so effective. This paper proposes such an account: positive preferences and negative constraints are structurally asymmetric. Positive preferences ("which is better") encode continuously coupled, context-dependent human values that cannot be exhaustively specified -- leading models to learn surface correlates such as agreement with the user (sycophancy). Negative constraints ("what is wrong") encode discrete, finite, independently verifiable prohibitions that can converge to a stable boundary. This asymmetry -- rooted in Popper's falsification logic and the epistemology of negative knowledge -- explains both the sycophancy failure of preference-based RLHF and the surprising effectiveness of negative-signal methods. We argue that alignment research should shift its center of gravity from "learning what humans prefer" to "learning what humans reject," and offer testable predictions for this framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。