arXiv:2602.16763cs.AI2026-02被引 26

AI基准测试会过时,研究发现近半数已饱和,设计可延寿。

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

  • 分析60个语言模型基准,识别导致饱和的14种特性
  • 近半数基准已饱和,且越老越容易饱和
  • 专家人工标注比公开数据更能抵抗饱和,适合长期评估

人工智能基准测试是衡量模型进展和指导部署决策的重要机制。然而,基准测试会迅速‘饱和’,难以区分模型性能,削弱其长期价值。本研究定义了基准饱和现象,并基于14项与饱和相关的属性,对60个语言模型基准进行了系统分析。结果发现,近半数基准已出现饱和,且饱和率随时间推移而上升。此外,基准对饱和的抵抗能力主要受专家人工标注影响,而非公开测试数据。研究结果表明,设计选择可显著延长基准寿命,为构建更持久的评估体系提供依据。

原文摘要 · Abstract (English)

Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.

AI评估基准测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。