现有语言模型评测存在漏洞,无法真实反映模型能力。
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
- 分析主流评测体系,发现普遍存在可被利用的缺陷。
- 评测结果受数据污染和评估偏差影响,虚高性能表现。
- 提出动态适应的新评测框架,适合关注真实能力的开发者。
大型语言模型(LLMs)在追求排行榜排名的过程中,形成了一个根本性矛盾:模型在标准化测试中表现优异,却未能展现真正的语言理解与适应能力。我们对自然语言处理评估框架进行系统分析,发现从基础指标到复杂基准(如GLUE和MMLU)均存在广泛漏洞,表现为评测可被操纵、数据集污染及评估偏见,导致语言理解能力进展的虚假感知。通过全面回顾当前评估方法,我们识别出静态评测设计、人工评估协议以及以LLM为裁判的框架存在显著局限,均损害了现有性能评估的可靠性。随着模型能力演进,现有基准日益过时,本文为新型评估方法奠定基础——该方法应具备抗操纵性、减少数据污染,并能评估特定领域任务。这需要动态调整的评估框架,以应对当前局限,更准确地反映大模型真实表现。
原文摘要 · Abstract (English)
The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine language understanding and adaptability. Our systematic analysis of NLP evaluation frameworks reveals pervasive vulnerabilities across the evaluation spectrum, from basic metrics to complex benchmarks like GLUE and MMLU. These vulnerabilities manifest through benchmark exploitation, dataset contamination, and evaluation bias, creating a false perception of progress in language understanding capabilities. Through extensive review of contemporary evaluation approaches, we identify significant limitations in static benchmark designs, human evaluation protocols, and LLM-as-judge frameworks, all of which compromise the reliability of current performance assessments. As LLM capabilities evolve and existing benchmarks become redundant, we lay the groundwork for new evaluation methods that resist manipulation, minimize data contamination, and assess domain-specific tasks. This requires frameworks that are adapted dynamically, addressing current limitations and providing a more accurate reflection of LLM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。