arXiv:2604.10673cs.AIcs.HC2026-04被引 1

AI对齐不只是执行原则,更需在具体情境中解读与权衡。

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment

  • 从解释学视角看,对齐本质是上下文相关的判断过程。
  • 实证发现多数偏好数据涉及原则冲突或无差异,无法唯一决定选择。
  • 模型部署时的行为分布才暴露真实对齐问题,静态测试易遗漏。

AI对齐常被理解为让系统遵循既定原则或人类偏好,但一般性原则很少能自行应用于具体情境。当原则冲突、过于宽泛或事实不清时,需额外的判断行为。本文从解释学角度分析这一判断环节,主张对齐包含解释成分:即需根据具体情境判断原则如何解读、应用与优先排序。我们联系近期实证研究,发现大量偏好标注数据属于原则冲突或无差异情形,原则集无法唯一确定决策。由此推导出操作后果:因这些判断体现在模型部署时的行为分布中,许多对齐相关选择仅在部署响应分布中显现。为形式化此点,我们区分部署诱发与语料诱发评估,指出当两种响应分布不同时,离策略审计可能无法捕捉关键对齐失效。

原文摘要 · Abstract (English)

AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles rarely determine their own application in concrete cases. When principles conflict, when they are too broad to settle a situation, or when the relevant facts are unclear, an additional act of judgment is required. This paper analyzes that step through the lens of hermeneutics and argues that alignment therefore includes an interpretive component: it involves context-sensitive judgments about how principles should be read, applied, and prioritized in practice. We connect this claim to recent empirical findings showing that a substantial portion of preference-labeling data falls into cases of principle conflict or indifference, where the principle set does not uniquely determine a decision. We then draw an operational consequence: because such judgments are expressed in behavior, many alignment-relevant choices appear only in the distribution of responses a model generates at deployment time. To formalize this point, we distinguish deployment-induced and corpus-induced evaluation and show that off-policy audits can fail to capture alignment-relevant failures when the two response distributions differ. We argue that principle-specified alignment includes a context-dependent interpretive component.

AI对齐解释学原则冲突部署评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。