强偏好会显著影响偏好模型的预测稳定性,威胁AI对齐安全。
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
- 分析偏好模型对强偏好变化的敏感性,揭示其内在脆弱性。
- 发现当偏好概率接近0或1时,其他偏好预测可能剧烈波动。
- 提醒开发者关注极端偏好对对齐系统鲁棒性的潜在风险。
价值对齐旨在确保大语言模型等人工智能体的行为符合人类价值观,是保障系统安全与可信的关键。其核心在于将人类偏好建模为人类价值观的体现。本文通过考察偏好模型对输入偏好概率微小变化的敏感性,研究价值对齐的鲁棒性。我们理论上分析了广泛使用的布拉德利-特瑞模型(Bradley-Terry)和普拉克特-卢斯模型(Plackett-Luce)的稳健性。结果表明,当某些偏好占主导地位(即概率接近0或1)时,其他偏好预测的概率可能发生显著变化。我们识别出该敏感性显著的特定条件,并讨论其对人工智能系统对齐鲁棒性与安全性的实际影响。
原文摘要 · Abstract (English)
Value alignment, which aims to ensure that large language models (LLMs) and other AI agents behave in accordance with human values, is critical for ensuring safety and trustworthiness of these systems. A key component of value alignment is the modeling of human preferences as a representation of human values. In this paper, we investigate the robustness of value alignment by examining the sensitivity of preference models. Specifically, we ask: how do changes in the probabilities of some preferences affect the predictions of these models for other preferences? To answer this question, we theoretically analyze the robustness of widely used preference models by examining their sensitivities to minor changes in preferences they model. Our findings reveal that, in the Bradley-Terry and the Placket-Luce model, the probability of a preference can change significantly as other preferences change, especially when these preferences are dominant (i.e., with probabilities near 0 or 1). We identify specific conditions where this sensitivity becomes significant for these models and discuss the practical implications for the robustness and safety of value alignment in AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。