arXiv:2606.31446cs.CLcs.CV2026-06

修复文档分类数据集的标签错误和重复,提升模型真实性能评估

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

论文配图:Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap
图 1 · 摘自论文原文
  • 系统检测并修正12%的标签错误,识别并清除约35%的测试训练重叠
  • 去除错误标签后分类准确率提升,但移除重复数据后准确率下降
  • 修正后的数据显著增强模型在分布外任务的泛化能力,提升8.1~14个百分点

RVL-CDIP 是文档分类模型常用的基准数据集,但存在大量标签错误和非显著的测试-训练重叠,可能误导模型性能评估。本文通过(1)发现并修正标签错误,(2)检测并处理测试-训练重叠,构建了多个改进版 RVL-CDIP。分析显示原数据集含12%标签错误和约35%测试-训练重复。修正标签错误后分类准确率提升,而移除重复样本后准确率下降。此外,在 RVL-CDIP-N(分布外基准)上评估发现,使用修正后的数据训练可显著提升模型泛化能力:监督模型平均准确率提升8.1个百分点,最高达14个百分点。

原文摘要 · Abstract (English)

RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overlap, both of which may impact model performance metrics. In this paper, we address these two problems by (1) finding and fixing label errors, and (2) detecting and addressing test-train overlap. We produce several variations of RVL-CDIP with label error and test-train overlap fixes, and benchmark document classification performance on these new RVL-CDIP variations. Our rigorous analysis of RVL-CDIP finds that the corpus contains 12\% label error and approximately 35% test-train duplication. Remediation sees improvements in classification accuracy when errors are removed, but sees decreases in accuracy when duplicates are removed. We additionally evaluate models on RVL-CDIP-N, an out-of-distribution benchmark, finding that training on error-corrected data substantially improves OOD generalization, with supervised models gaining an average of 8.1 percentage points in accuracy and improvements as large as 14 percentage points.

数据集清洗文档分类模型评估泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。