提出用法律中的裁量权概念分析AI对齐中的判断自由度,揭示隐性偏差与模型模仿风险。
AI Alignment at Your Discretion
- 引入法律裁量权理论,量化标注者在对齐判断中的自主空间。
- 发现人类与算法在安全对齐标注中存在显著裁量差异,且模型会自发形成裁量模式。
- 揭示当前对齐流程中未被重视的裁量层,警示其引发的偏见与原则失效风险。
在人工智能对齐中,标注者(无论是人还是算法)需拥有充分裁量权来判断哪些模型输出更优或更安全,这种权力称为对齐裁量。然而,裁量权至今未被深入研究,带来两大风险:(i) 标注者可能任意行使裁量权;(ii) 模型可能无法有效模仿该裁量。为此,本文借鉴法律中关于裁量权的理论,分析当对齐原则冲突或模糊时,决策权如何分配与行使。我们提出一套系统性指标,用于分析对齐裁量的出现时机与方式,从而同时观察上述两类风险。进一步区分人类与算法裁量,并在多个安全对齐数据集上测量两者差异。结果揭示了以往未被察觉的对齐过程中的多层裁量机制。此外,我们证明基于这些数据集训练的算法会发展出自身形式的裁量行为,进而挑战对齐原则存在的意义。本文是首次尝试正式化当前对齐流程中的这一核心缺口,呼吁社区加强对对齐裁量的审视与控制。
原文摘要 · Abstract (English)
In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely unexamined, posing two risks: (i) annotators may use their power of discretion arbitrarily, and (ii) models may fail to mimic this discretion. To study this phenomenon, we draw on legal concepts of discretion that structure how decision-making authority is conferred and exercised, particularly in cases where principles conflict or their application is unclear or irrelevant. Extended to AI alignment, discretion is required when alignment principles and rules are (inevitably) conflicting or indecisive. We present a set of metrics to systematically analyze when and how discretion in AI alignment is exercised, such that both risks (i) and (ii) can be observed. Moreover, we distinguish between human and algorithmic discretion and analyze the discrepancy between them. By measuring both human and algorithmic discretion over safety alignment datasets, we reveal layers of discretion in the alignment process that were previously unaccounted for. Furthermore, we demonstrate how algorithms trained on these datasets develop their own forms of discretion in interpreting and applying these principles, which challenges the purpose of having any principles at all. Our paper presents the first step towards formalizing this core gap in current alignment processes, and we call on the community to further scrutinize and control alignment discretion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。