arXiv:2410.09247cs.LGcs.AI2024-10被引 15

发现大模型评测分数虚高,通过回溯构造新测试集揭示真实性能差距

Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts

  • 回溯构建与原数据统计上不可区分的测试集
  • 在新测试集上部分模型得分低16个百分点
  • 提醒研究者警惕评测数据泄露问题

许多大语言模型的训练数据中混入了公开测试数据,导致现有评测基准被污染,无法真实反映模型能力。为解决此问题,本文提出系统方法:(1)为指定数据集回溯构建保留集;(2)证明该回溯保留集与原数据在统计上不可区分;(3)对比模型在原始集与保留集上的表现,量化因数据公开造成的性能虚高。以TruthfulQA为例,我们构建并发布Retro-Misconceptions,评估20个LLM,发现部分模型得分虚高达16个百分点。结果表明,公开评测分数未必真实反映模型性能,凸显改进数据管理实践的重要性。

原文摘要 · Abstract (English)

The training data for many Large Language Models (LLMs) is contaminated with test data. This means that public benchmarks used to assess LLMs are compromised, suggesting a performance gap between benchmark scores and actual capabilities. Ideally, a private holdout set could be used to accurately verify scores. Unfortunately, such datasets do not exist for most benchmarks, and post-hoc construction of sufficiently similar datasets is non-trivial. To address these issues, we introduce a systematic methodology for (i) retrospectively constructing a holdout dataset for a target dataset, (ii) demonstrating the statistical indistinguishability of this retro-holdout dataset, and (iii) comparing LLMs on the two datasets to quantify the performance gap due to the dataset's public availability. Applying these methods to TruthfulQA, we construct and release Retro-Misconceptions, on which we evaluate twenty LLMs and find that some have inflated scores by as much as 16 percentage points. Our results demonstrate that public benchmark scores do not always accurately assess model properties, and underscore the importance of improved data practices in the field.

大模型评测数据泄露性能虚高

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。