发现大模型性能报告被高估,因标签错误掩盖真实能力
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
- 用多个大模型作为评委,自动检测数据集中的标签错误
- 修正错误标签后,模型性能显著提升,部分表现翻倍
- 适合关注模型真实能力、数据质量的研究者与开发者
NLP基准依赖标准化数据集进行训练与评估,对领域进步至关重要。传统专家标注虽质量高,但成本难以随现代模型需求扩展;众包标注虽可扩展,却常牺牲精度与一致性。近年来大语言模型(LLMs)为改进标注流程提供了新可能,尤其在检测现有数据集中的标签错误方面。本文采用LLM-as-a-judge方法,通过集成多个大模型识别潜在误标样本,对TRUE基准中的四个事实一致性数据集及SummEval(多维度摘要质量评分)进行了案例研究。我们实证分析了专家、众包与基于LLM标注的协议度、标签质量与效率,揭示了大量标签错误。修正后,模型报告性能出现显著提升,表明许多所谓‘模型失败’实为标签错误所致。同时讨论了误标数据的影响,并提出训练阶段缓解策略以提升性能。
原文摘要 · Abstract (English)
NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field. Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models. While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency. Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets. In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples. We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likert-scale ratings of summary quality across multiple dimensions. We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method. Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance. This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures. Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。