arXiv:2605.28802cs.CL2026-05中稿 · EMNLP

让AI学习标注者独特推理逻辑,提升解释一致性。

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

论文配图:Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
图 1 · 摘自论文原文
  • 通过跨标注者偏好优化,捕捉个体标注习惯。
  • 在两个任务中验证,标注者行为模式可被有效学习。
  • 适合需要可解释标注的AI训练与质量评估场景。

自由文本解释将人类标注差异(HLV)从标签分歧扩展至标注者决策背后的推理与偏好。我们研究大语言模型(LLMs)能否学习并复现这种标注者特异性标签-解释行为。在自然语言推理和释义判断两个句子对任务中,各包含四位标注者。首先分析标注者是否表现出稳定的个体模式,发现单次标注层面因输入内容影响强而模式弱,但经输入内容缩减与标注者层级聚合后,模式变得可检测。接着对比提示(prompting)与监督微调(SFT)基线,提出跨标注者偏好优化(CAPO),通过对比目标标注者的回答与其他有效但非目标特定的标注,强化目标特异性响应。实验表明,提示方法受限且不稳定,SFT能更好捕捉标注者特异性行为,而CAPO进一步提升了聚合感知模仿与基于判断的归因能力,同时在人工验证下保持目标特异性推理模式。结果表明,HLV可被学习为标注者特异性标签-解释行为,为基于标注历史而非仅标签的可扩展解释式标注提供了路径。

原文摘要 · Abstract (English)

Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large language models (LLMs) can learn and reproduce such annotator-specific label-explanation behavior. Using two sentence-pair tasks with four annotators each -- natural language inference and paraphrase judgment -- we first analyze whether annotators exhibit stable individual patterns. We find that such patterns are weak at the single-annotation level due to strong input-content effects, but become detectable after input-content reduction and annotator-level aggregation. We then compare prompting and supervised fine-tuning (SFT) baselines and propose cross-annotator preference optimization (CAPO), which contrasts a target annotator's response with other valid but less target-specific annotations for the same input. Experiments show that prompting is limited and unstable, SFT better captures annotator-specific behavior, and CAPO further improves aggregation-aware imitation and judge-based attribution while preserving target-specific reasoning patterns under human validation. Overall, our results show that HLV can be learned as annotator-specific label-explanation behavior, suggesting a path toward scalable explanation-based annotation grounded in annotator histories rather than labels alone.

大模型标注一致性解释生成偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。