分析上千篇论文发现人类评估存在严重报告缺失,影响结果可信度。
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

- 系统审查284篇论文+1800+篇半自动分析,梳理评估协议透明度
- 73%研究未完整报告评估者数量与标注标准,关键信息缺失
- 建议建立20项可报告标准,助力未来研究可复现
人类评估在生成文本质量评价中至关重要,但其可靠性依赖于透明且完整的评估协议——而当前实践中这类细节常被忽略。本文对2023至2025年CL会议发表的长文本生成相关论文开展大规模分析,人工审查284篇论文,并对1800余篇采用LLM辅助分析。我们定义了20项与可复现性相关的报告标准,系统考察社区中的报告规范。结果发现,人类评估研究普遍存在重要设计信息报告不全的问题,导致对测量内容、评估方式、评估者构成及结果解释存在模糊性。基于此,我们提出具体改进建议,以推动未来研究实现更高透明度与可复现性。分析代码与标注数据集已开源:https://github.com/larchlab/Illusions-of-the-Gold-Standard
原文摘要 · Abstract (English)
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。