从静态到动态评估,应对大模型数据污染风险
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
- 提出动态基准测试设计原则,解决评估标准缺失问题
- 揭示现有动态基准的局限性,推动评测体系规范化
- 适合关注大模型评测可信性的研究者与开发者
大语言模型依赖海量互联网训练语料,数据污染问题日益凸显。为降低数据泄露风险,模型评测正从静态基准向动态基准演进。本文深入分析现有静态与动态评测方法,指出静态方法存在固有缺陷,并揭示动态基准缺乏统一评估标准的关键短板。基于此,提出一系列最优动态评测设计原则,系统评估现有动态基准的不足。本综述全面梳理了数据污染研究的最新进展,为未来研究提供清晰指引。相关方法持续收集于开源GitHub仓库,可供持续追踪。
原文摘要 · Abstract (English)
Data contamination has received increasing attention in the era of large language models (LLMs) due to their reliance on vast Internet-derived training corpora. To mitigate the risk of potential data contamination, LLM benchmarking has undergone a transformation from static to dynamic benchmarking. In this work, we conduct an in-depth analysis of existing static to dynamic benchmarking methods aimed at reducing data contamination risks. We first examine methods that enhance static benchmarks and identify their inherent limitations. We then highlight a critical gap-the lack of standardized criteria for evaluating dynamic benchmarks. Based on this observation, we propose a series of optimal design principles for dynamic benchmarking and analyze the limitations of existing dynamic benchmarks. This survey provides a concise yet comprehensive overview of recent advancements in data contamination research, offering valuable insights and a clear guide for future research efforts. We maintain a GitHub repository to continuously collect both static and dynamic benchmarking methods for LLMs. The repository can be found at this link.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。