首次大规模审计NLP论文中人工标注报告,揭示透明度不足问题。
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

- 构建统一标注报告分类体系,用LLM自动提取2667个任务
- 仅41%报告训练细节,68%未提一致性指标,模型评估更不透明
- 适合关注可复现性与数据可信度的研究者参考
人工标注是自然语言处理研究的实证基础,但多数论文未明确标注人员构成及过程控制。本文对2018至2025年主要NLP会议论文进行首次大规模、任务级审计,分析标注信息报告的完整性与演变趋势。提出统一标注报告分类体系,并通过基于LLM的抽取管道在Annotated-gold(41篇论文,72个任务)上验证,最优模型与人类标注者达成相近一致性(Krippendorff's alpha=0.606 vs 0.585)。利用该管道构建Annotated-llm,涵盖1603篇ACL系列论文中的2667个标注任务。发现尽管招募策略、专家背景等操作细节常被报告,但训练流程、语言能力、报酬、社会人口学特征、仲裁机制及一致率等关键信息仍普遍缺失,尤其在模型评估类研究中。结果显示,尽管整体报告质量有所提升,仍存在显著不均衡,本文提出可扩展框架与最低报告标准,以增强人工标注的可靠性、可复现性与可解释性。
原文摘要 · Abstract (English)
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled. We provide the first large-scale, task-level audit of human annotation reporting across major NLP venues, asking which annotation details are documented, which are missing, and how reporting varies across time, topic, venue, and intended use of human judgment. We introduce a unified taxonomy of annotation-reporting practices and validate an LLM-assisted extraction pipeline against Annotated-gold, a human-adjudicated gold standard of 41 papers and 72 annotation tasks, where the best model reaches human-comparable agreement with adjudicated labels, with Krippendorff's alpha of 0.606 versus 0.585 for human-human agreement. Using this pipeline, we construct Annotated-llm, a dataset covering ACL-venue papers from 2018-2025, with 2,667 extracted annotation tasks from 1,603 papers, and find that papers frequently report operational details such as recruitment strategies, annotator expertise, and annotation volume, but often omit details needed to assess annotation validity, including training, language proficiency, compensation, socio-demographics, adjudication, and agreement values, especially in model-evaluation studies. Our results show that annotation reporting in NLP has improved over time but remains uneven, and they establish a scalable framework and bare-minimum reporting recommendations for making human annotation more reliable, reproducible, and interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。