用AI推理引导人工标注,提升一致性且减少修改。
ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation
- 让人类标注者先独立判断,再看AI推理后修订
- 标注一致性和共识度显著提升,仅少量标签被修改
- 适合需要高一致性标注的NLP评估场景
人工标注是自然语言处理评估的核心,但主观任务常因标注者差异导致结果不一。尽管大语言模型(LLMs)能提供结构化推理辅助标注,其对人类标注行为的影响仍不明确。本文提出「ReasonScaffold」——一种基于推理引导的标注协议,展示模型生成的解释但隐藏预测标签。通过受控实验,采用类似德尔菲法的双阶段流程:标注者先独立标注,再在看到模型推理后修订判断。在情感分类与观点检测任务上评估该方法,分析标注者间一致性及修订行为变化。引入标注者努力代理(AEP)量化因推理引发的标签修改比例。结果显示,接触推理后标注一致性提高,但修订比例极低,表明推理有助于解决模糊案例,而不会引发大规模重标。研究揭示了推理解释如何提升标注一致性,并证明推理引导是一种实用的人机协同标注机制。
原文摘要 · Abstract (English)
Human annotation is central to NLP evaluation, yet subjective tasks often exhibit substantial variability across annotators. While large language models (LLMs) can provide structured reasoning to support annotation, their influence on human annotation behavior remains underexplored. We introduce \textbf{ReasonScaffold}, a scaffolded reasoning annotation protocol that exposes LLM-generated explanations while withholding predicted labels. We study how reasoning affects human annotation behavior in a controlled setting, rather than evaluating annotation accuracy. Using a two-pass protocol inspired by Delphi-style revision, annotators first label instances independently and then revise their decisions after viewing model-generated reasoning. We evaluate the approach on sentiment classification and opinion detection tasks, analyzing changes in inter-annotator agreement and revision behavior. To quantify these effects, we introduce the Annotator Effort Proxy (AEP), a metric capturing the proportion of labels revised after exposure to reasoning. Our results show that exposure to reasoning is associated with increased agreement, along with minimal revision, suggesting that reasoning helps resolve ambiguous cases without inducing widespread changes. These findings provide insight into how reasoning explanations shape annotation consistency and highlight reasoning-based scaffolds as a practical mechanism for human--AI co-annotation workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。