自动进化评测集,让模型测试更公平真实
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- 用多模型竞争生成新题,保留原测试目标
- 可发现新漏洞,提升难度且保持测试一致性
- 适合关注模型真实能力评估的研究者
评测集是衡量大模型能力、指导模型发展的核心工具,但预训练数据泄露导致模型可能仅靠记忆而非泛化得分,从而虚高评分、扭曲对比并误导进展判断。我们提出 ArenaBencher,一个与模型无关的自动评测集演化框架,在不破坏可比性的前提下更新测试题。给定现有评测集和多样模型池,该框架推断每道题的核心能力,生成符合原目标的候选问答对,利用大模型作为裁判验证正确性与意图,并聚合多模型反馈,选择能暴露共性弱点的题目。过程通过上下文示范迭代引导生成更具挑战性和诊断性的题目。我们在数学求解、常识推理和安全领域应用该方法,证明其能产出经验证、多样化、公平的更新内容,揭示新失败模式,提升难度同时保持目标一致,增强模型区分度。该框架为随基础模型快速演进而持续演化评测集提供了可扩展路径。
原文摘要 · Abstract (English)
Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their validity. Models can match memorized content rather than demonstrate true generalization, which inflates scores, distorts cross-model comparisons, and misrepresents progress. We introduce ArenaBencher, a model-agnostic framework for automatic benchmark evolution that updates test cases while preserving comparability. Given an existing benchmark and a diverse pool of models to be evaluated, ArenaBencher infers the core ability of each test case, generates candidate question-answer pairs that preserve the original objective, verifies correctness and intent with an LLM as a judge, and aggregates feedback from multiple models to select candidates that expose shared weaknesses. The process runs iteratively with in-context demonstrations that steer generation toward more challenging and diagnostic cases. We apply ArenaBencher to math problem solving, commonsense reasoning, and safety domains and show that it produces verified, diverse, and fair updates that uncover new failure modes, increase difficulty while preserving test objective alignment, and improve model separability. The framework provides a scalable path to continuously evolve benchmarks in step with the rapid progress of foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。