arXiv:2505.14633cs.CLcs.AI2025-05被引 18

通过测试AI的价值优先级,提前发现其潜在风险行为。

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

  • 构建价值评估流水线LitmusValues,量化AI对多种价值的偏好。
  • 在多类安全困境中,模型价值排序可预测已知与未知风险行为。
  • 适用于评估强AI模型的安全性,尤其关注权力追求等高风险场景。

随着更强模型的出现,检测其风险行为愈发困难,部分模型甚至采用对齐伪装(Alignment Faking)来规避检测。受人类危险行为常由强烈价值观驱动的启发,我们提出:识别AI内部的价值取向可作为早期预警机制。为此,我们设计了LitmusValues评估流水线,用于揭示模型在多种价值类别上的优先级。进一步,我们构建了AIRiskDilemmas数据集,包含一系列涉及价值冲突的、与AI安全风险相关的难题(如权力寻求)。通过分析模型在这些难题中的整体选择,我们获得一组自洽的价值优先级预测,成功揭示潜在风险。结果表明,即便看似无害的价值(如关怀Care),也能有效预测在已有难题和未见风险任务(HarmBench)中的风险行为。

原文摘要 · Abstract (English)

Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activities that may hurt others) are sometimes guided by strongly-held values, we believe that identifying values within AI models can be an early warning system for AI's risky behaviors. We create LitmusValues, an evaluation pipeline to reveal AI models' priorities on a range of AI value classes. Then, we collect AIRiskDilemmas, a diverse collection of dilemmas that pit values against one another in scenarios relevant to AI safety risks such as Power Seeking. By measuring an AI model's value prioritization using its aggregate choices, we obtain a self-consistent set of predicted value priorities that uncover potential risks. We show that values in LitmusValues (including seemingly innocuous ones like Care) can predict for both seen risky behaviors in AIRiskDilemmas and unseen risky behaviors in HarmBench.

AI安全价值对齐风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。