arXiv:2507.11405cs.CL2025-07EMNLP被引 14

提出量化大模型评估中数据污染风险的新方法,提升评测可信度。

DCR: Quantifying Data Contamination in LLMs Evaluation

  • 通过语义、信息、数据和标签四层分析,构建可解释的污染检测框架。
  • 在9个不同规模模型上验证,调整后准确率误差控制在4%以内。
  • 轻量高效且透明,适合日常评测中纳入污染评估环节。

大语言模型(LLMs)的快速发展引发了对基准数据污染(BDC)的担忧,即模型在训练过程中无意记忆评估数据,导致性能指标虚高,削弱了真实泛化能力的评估。本文提出数据污染风险(DCR)框架,一种轻量级、可解释的检测与量化管道,可在语义、信息、数据和标签四个粒度层级识别并评估污染风险。通过模糊推理系统融合各层污染得分,生成统一的DCR因子,用于修正原始准确率以反映污染感知性能。在9个规模从0.5B到72B的LLMs上,针对情感分析、假新闻检测和算术推理任务进行验证,使用DCR因子调整后的准确率与无污染基线相比平均误差不超过4%。该框架强调计算效率与透明性,为常规评估中融入污染评估提供了实用工具,有助于实现更公平的模型比较与提升评测可信度。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has heightened concerns about benchmark data contamination (BDC), where models inadvertently memorize evaluation data during the training process, inflating performance metrics, and undermining genuine generalization assessment. This paper introduces the Data Contamination Risk (DCR) framework, a lightweight, interpretable pipeline designed to detect and quantify BDC risk across four granular levels: semantic, informational, data, and label. By synthesizing contamination scores via a fuzzy inference system, DCR produces a unified DCR Factor that adjusts raw accuracy to reflect contamination-aware performance. Validated on 9 LLMs (0.5B-72B) across sentiment analysis, fake news detection, and arithmetic reasoning tasks, the DCR framework reliably diagnoses contamination severity and with accuracy adjusted using the DCR Factor to within 4% average error across the three benchmarks compared to the uncontaminated baseline. Emphasizing computational efficiency and transparency, DCR provides a practical tool for integrating contamination assessment into routine evaluations, fostering fairer comparisons and enhancing the credibility of LLM benchmarking practices.

大模型评估数据污染可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。