arXiv:2501.18771cs.CLcs.AI2025-01ICML被引 16

发现大模型翻译评估因数据污染虚高,8B模型误差比1B高出2.5倍。

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

  • 系统测试不同污染程度对1B与8B模型的影响
  • 源目标双污染使BLEU分最高提升30点,8B模型增幅是1B的2.5倍
  • 污染时间分布和频次影响评估偏差,资源少语言更敏感

数据污染——即评估样本意外出现在预训练数据中——会破坏评估基准的有效性。本文针对10亿和80亿参数规模的语言模型,在机器翻译任务上开展严谨分析。从精心去污的训练-测试划分出发,系统引入不同阶段、规模和格式的污染,以隔离其影响并量化性能指标变化。实验表明,源语言与目标语言同时污染会使BLEU分数显著虚高,且80亿模型的膨胀效应是10亿模型的2.5倍(最高达30个BLEU点)。而仅源或仅目标污染产生的过估计较小且不一致。最后,我们研究了污染样本的时间分布与频率如何影响不同语言资源条件下评估结果的虚高程度。

原文摘要 · Abstract (English)

Data contamination -- the accidental consumption of evaluation examples within the pre-training data -- can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources.

大模型评估数据污染机器翻译BLEU

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。