arXiv:2601.11920cs.CLcs.AI2026-01被引 2

通过分解错误类型,提升大模型在主观标注中的可靠性。

Enhancing LLM-Based Data Annotation with Error Decomposition

  • 将错误分为模型自身与任务固有两类,细分为边界模糊和概念误判。
  • 在4个教育标注任务中验证,发现高对齐率不等于高质量标注。
  • 适合需精准标注的心理学、教育类研究,帮助优化模型使用策略。

大语言模型为数据标注提供了可扩展的替代方案,但在涉及心理构念等主观任务上表现不一,易出错。现有评估常将所有错误合并为单一对齐指标,掩盖了不同错误对分析结果的差异化影响。本文提出一种诊断性评估范式,引入人工参与,区分任务固有模糊性与模型驱动误差,并在序数标注任务中细化该方法:(1)构建双维度错误分类体系,按来源(模型特有或任务固有)与类型(边界模糊或概念误判)分类;(2)设计轻量级人工标注测试,估算任务固有模糊性;(3)提出计算方法,依据分类体系分解观测到的标注错误。在四个教育标注任务中验证,证明该范式兼具理论合理性与实用价值。理论上,揭示了特定任务中过高的对齐率不现实,且单一对齐指标无法准确反映标注质量;实践中,可作为低成本诊断工具,评估任务是否适合用大模型标注,并提供技术优化方向。

原文摘要 · Abstract (English)

Large language models offer a scalable alternative to human coding for data annotation tasks, enabling the scale-up of research across data-intensive domains. While LLMs are already achieving near-human accuracy on objective annotation tasks, their performance on subjective annotation tasks, such as those involving psychological constructs, is less consistent and more prone to errors. Standard evaluation practices typically collapse all annotation errors into a single alignment metric, but this simplified approach may obscure different kinds of errors that affect final analytical conclusions in different ways. Here, we propose a diagnostic evaluation paradigm that incorporates a human-in-the-loop step to separate task-inherent ambiguity from model-driven inaccuracies and assess annotation quality in terms of their potential downstream impacts. We refine this paradigm on ordinal annotation tasks, which are common in subjective annotation. The refined paradigm includes: (1) a diagnostic taxonomy that categorizes LLM annotation errors along two dimensions: source (model-specific vs. task-inherent) and type (boundary ambiguity vs. conceptual misidentification); (2) a lightweight human annotation test to estimate task-inherent ambiguity from LLM annotations; and (3) a computational method to decompose observed LLM annotation errors following our taxonomy. We validate this paradigm on four educational annotation tasks, demonstrating both its conceptual validity and practical utility. Theoretically, our work provides empirical evidence for why excessively high alignment is unrealistic in specific annotation tasks and why single alignment metrics inadequately reflect the quality of LLM annotations. In practice, our paradigm can be a low-cost diagnostic tool that assesses the suitability of a given task for LLM annotation and provides actionable insights for further technical optimization.

大模型标注质量错误分解主观标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。