arXiv:2412.13670cs.CLcs.LG2024-12ACL被引 47

自动构建无数据泄露的评测基准,解决大模型评估中的训练数据污染问题。

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge

  • 用模型训练集外的新知识构造测试样本,确保无数据泄露。
  • 实验发现模型截止时间前已存在数据污染,新基准有效规避该问题。
  • 全流程自动化,无需人工参与,显著降低评测维护成本。

数据污染会因测试数据被纳入新模型训练集而干扰大模型的公平评估。现有方法通过引入新采集数据更新基准,但无法保证无污染,且依赖大量人工。本文提出 AntiLeakBench,一种全自动抗泄露评测框架。不直接使用新数据,而是显式构造模型训练集之外的新知识样本,确保严格无污染评估。设计完全自动化工作流,实现基准的自动构建与更新,无需人工干预,极大降低维护成本以适配新兴大模型。大量实验表明,数据污染可能存在于模型截止时间之前,而 AntiLeakBench 有效克服此挑战。

原文摘要 · Abstract (English)

Data contamination hinders fair LLM evaluation by introducing test data into newer models' training sets. Existing studies solve this challenge by updating benchmarks with newly collected data. However, they fail to guarantee contamination-free evaluation as the newly collected data may contain pre-existing knowledge, and their benchmark updates rely on intensive human labor. To address these issues, we in this paper propose AntiLeak-Bench, an automated anti-leakage benchmarking framework. Instead of simply using newly collected data, we construct samples with explicitly new knowledge absent from LLMs' training sets, which thus ensures strictly contamination-free evaluation. We further design a fully automated workflow to build and update our benchmark without human labor. This significantly reduces the cost of benchmark maintenance to accommodate emerging LLMs. Through extensive experiments, we highlight that data contamination likely exists before LLMs' cutoff time and demonstrate AntiLeak-Bench effectively overcomes this challenge.

大模型评估数据污染自动化基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。