arXiv:2507.02593cs.CLcs.HC2025-07被引 6

重新思考人类标注差异下的主动学习,让模型更懂真实标注的复杂性。

Revisiting Active Learning under (Human) Label Variation

  • 将标注差异分解为人类标注变异性信号与错误噪声,改进主动学习假设。
  • 提出贯穿主动学习全流程的HLV感知框架,支持选例、选人和标签表示优化。
  • 首次系统讨论大模型作为标注者在含差异场景中的应用潜力。

高质量标注数据仍是实际监督学习的瓶颈。尽管标注差异(LV)在自然语言处理中普遍存在,但标注框架仍常假设单一真实标签。这忽略了人类标注变异性(HLV)——即对同一实例存在合理分歧——作为信息信号的价值。主动学习(AL)虽能优化有限标注预算,但其常用简化假设在承认HLV时往往失效。本文重新审视关于真实性的基础假设,强调需将观测到的标注差异拆分为信号(如HLV)与噪声(如标注错误)。我们综述了AL与(H)LV领域对此区分的处理或忽视,并提出一个包含实例选择、标注者选择与标签表示的全过程HLV感知框架。进一步探讨大语言模型作为标注者的整合可能性。本工作旨在为面向真实标注复杂性的主动学习建立概念基础。

原文摘要 · Abstract (English)

Access to high-quality labeled data remains a limiting factor in applied supervised learning. While label variation (LV), i.e., differing labels for the same instance, is common, especially in natural language processing, annotation frameworks often still rest on the assumption of a single ground truth. This overlooks human label variation (HLV), the occurrence of plausible differences in annotations, as an informative signal. Similarly, active learning (AL), a popular approach to optimizing the use of limited annotation budgets in training ML models, often relies on at least one of several simplifying assumptions, which rarely hold in practice when acknowledging HLV. In this paper, we examine foundational assumptions about truth and label nature, highlighting the need to decompose observed LV into signal (e.g., HLV) and noise (e.g., annotation error). We survey how the AL and (H)LV communities have addressed -- or neglected -- these distinctions and propose a conceptual framework for incorporating HLV throughout the AL loop, including instance selection, annotator choice, and label representation. We further discuss the integration of large language models (LLM) as annotators. Our work aims to lay a conceptual foundation for HLV-aware active learning, better reflecting the complexities of real-world annotation.

主动学习标注差异人类标注大模型标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。