arXiv:2605.05329cs.AIcs.LG2026-05

用可解释模型分析标注者安全政策,看清分歧根源

Understanding Annotator Safety Policy with Interpretability

论文配图:Understanding Annotator Safety Policy with Interpretability
图 1 · 摘自论文原文
  • 通过标注行为建模标注者内在安全政策,无需额外标注
  • 准确预测反事实修改结果(>80%准确率),还原政策差异
  • 发现政策模糊与价值观多元,助力更透明、包容的安全设计

安全政策定义了人工智能输出的可接受与不可接受标准,指导数据标注与模型开发。然而标注分歧普遍存在,可能源于操作失误(标注者误解任务)、政策模糊(政策表述存在歧义)或价值多元(不同标注者对安全理解不同)。区分这些原因至关重要:操作失误需质量控制,模糊性需政策澄清,多元性则需纳入不同视角的讨论。但理解分歧成因困难。直接询问标注者理由成本高且不可靠,尤其对人类和大模型标注者而言,自述推理常不能反映真实决策过程。本文提出标注者政策模型(APM),一种可解释模型,仅从标注行为中学习标注者的内在安全政策,使标注者推理可见且可比,无需额外标注工作。验证显示APM能准确建模标注者安全政策(>80%准确率),可靠预测反事实编辑后的响应,并在受控环境中复现已知政策差异。应用于大模型与人类标注,揭示两大核心应用:(1) 通过展示标注者对安全指令的不同理解,暴露政策模糊;(2) 通过发现不同人口群体间系统性的安全优先级差异,揭示价值多元。这些能力共同支持更精准、透明和包容的安全政策设计。

原文摘要 · Abstract (English)

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational failures (annotators misunderstand or misexecute the task), policy ambiguity (policy wording leaves room for interpretation), or value pluralism (different annotators hold different perspectives on safety). Distinguishing these sources matters. For example, operational failures call for quality control, ambiguity calls for policy clarification, and pluralism calls for deliberation about incorporating diverse perspectives. Yet understanding why annotators disagree is difficult. Directly asking annotators for their reasoning is costly, substantially increasing annotation burden, and can be unreliable for both human and LLM annotators as self-reported reasoning often fails to reflect actual decision processes. We introduce Annotator Policy Models (APMs), interpretable models that learn annotators' internal safety policies from labeling behavior alone, making annotator reasoning visible and comparable without additional annotation effort. We validate that APMs accurately model annotator safety policy (>80% accuracy), faithfully predict responses to counterfactual edits, and recover known policy differences in controlled settings. Applying APMs to LLM and human annotations, we demonstrate two core applications: (1) surfacing policy ambiguity by revealing how annotators interpret safety instructions differently, and (2) surfacing value pluralism by uncovering systematic differences in safety priorities across demographic groups. Together, these capabilities support more targeted, transparent, and inclusive safety policy design.

可解释性安全策略标注分歧政策建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。