现有对齐方法忽视复杂情境下的失败,论文提出新框架让难题可见可治。
AI Alignment Breaks at the Edge
- 构建边缘情境诊断集,识别标准评估忽略的对齐失败
- 在91个边缘案例中,主流模型帮助性与安全性评分虚高
- 适合关注模型伦理、治理与复杂决策的研究者
通用对齐虽提升了平均情况下的帮助性与安全性,但当前对齐实践仍偏好评价自信且单轮响应。问题不仅在于模型在边缘情况失效,更在于评估机制使这些失败难以察觉。我们主张对齐必须超越平均情况评估,显化并可操作处理价值冲突、多方利益分歧与认知模糊情境中的失败。标量奖励将多元价值压缩为单一数值;数据与评估体系会过滤或遗漏对齐最困难的场景;治理缺乏解决争议的机制。这些盲区导致价值扁平化、表征丢失与不确定性盲视。本文提出‘边缘对齐’概念,作为检测、评估与治理的综合议程,旨在揭示这些问题并对接干预措施。边缘对齐并非单一训练目标,而是定义标准对齐应让位于保留多维价值结构、体现多元视角、支持不确定性交互的条件。一个包含91个边缘案例的试点集和四个当代模型的实验表明,常规帮助性与安全性评估可能掩盖过程性失败,而边缘感知评估能暴露这些问题。文章还提出操作性边缘信号、过程感知评估标准及三阶段流程栈,将对齐重构为动态规范治理的生命周期问题。
原文摘要 · Abstract (English)
General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. The problem is not only that models fail on edge cases; it is that current evaluation makes many of these failures hard to see. We take the position that alignment must move beyond average-case evaluation by making failures under value conflict, plural stakeholder disagreement, and epistemic ambiguity visible and actionable. Scalar rewards compress diverse values into a single number; data and evaluation regimes collapse, filter, or fail to elicit the cases where alignment is hardest; and governance often lacks mechanisms for adjudicating contested cases. These blind spots produce value flattening, representation loss, and uncertainty blindness. We use Edge alignment to name a detection, evaluation, and governance agenda for surfacing these failures and connecting them to appropriate interventions. Rather than a single training objective, Edge alignment defines the conditions under which standard alignment should yield to mechanisms that preserve multidimensional value structure, represent plural perspectives, and support uncertainty-aware interaction. A pilot diagnostic set of 91 edge cases and four contemporary models illustrates that ordinary helpfulness and safety readings can miss process failures that edge-aware evaluation exposes. We outline operational edge signals, process-aware evaluation criteria, and a three-phase process stack that reframes alignment as a lifecycle problem of dynamic normative governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。