arXiv:2508.05452cs.CL2025-08ACL被引 2

用动态测试集避免模型作弊,真实评估大模型能力

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models

  • 构建22万道研究生级题目库,每次评测随机抽题防数据泄露
  • 90%评分一致率,相对排名系统让模型性能对比更公平
  • 30个月追踪近60个模型,发现知识记忆有天花板

现有大语言模型评估依赖静态基准,易受数据污染和排行榜过拟合影响,掩盖真实能力。为此,我们提出LLMEval-Fair,一个动态评估框架。该框架基于自有的22万道研究生级别问题库,每次评估动态抽取未见测试集。其自动化流程通过抗污染数据整理、新型防作弊架构及校准的LLM评判机制(与人类专家达成90%一致),结合相对排名系统实现公平比较。对近60个领先模型开展为期30个月的纵向研究,揭示知识记忆存在性能上限,并暴露静态基准无法检测的数据污染漏洞。框架在排名稳定性与一致性上表现卓越,为动态评估范式提供有力实证。本研究为超越排行榜分数的真实能力评估提供可靠方法,推动可信评估标准发展。代码与数据公开于https://github.com/llmeval/LLMEval-Fair。

原文摘要 · Abstract (English)

Existing evaluation of Large Language Models (LLMs) on static benchmarks is vulnerable to data contamination and leaderboard overfitting, critical issues that obscure true model capabilities. To address this, we introduce LLMEval-Fair, a framework for dynamic evaluation of LLMs. LLMEval-Fair is built on a proprietary bank of 220k graduate-level questions, from which it dynamically samples unseen test sets for each evaluation run. Its automated pipeline ensures integrity via contamination-resistant data curation, a novel anti-cheating architecture, and a calibrated LLM-as-a-judge process achieving 90% agreement with human experts, complemented by a relative ranking system for fair comparison. A 30-month longitudinal study of nearly 60 leading models reveals a performance ceiling on knowledge memorization and exposes data contamination vulnerabilities undetectable by static benchmarks. The framework demonstrates exceptional robustness in ranking stability and consistency, providing strong empirical validation for the dynamic evaluation paradigm. LLMEval-Fair offers a robust and credible methodology for assessing the true capabilities of LLMs beyond leaderboard scores, promoting the development of more trustworthy evaluation standards. Our code and data are publicly available at https://github.com/llmeval/LLMEval-Fair.

大模型评估动态测试公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。