arXiv:2510.07575cs.AIcs.LG2025-10NeurIPS被引 8

AI评估乱象频发,需建立可信的统一评测体系。

Benchmarking is Broken -- Don't Let AI be its Own Judge

  • 提出由社区治理的实时评测框架,杜绝数据污染和选择性报告。
  • 通过密封执行与滚动更新题库,确保评测结果真实可靠。
  • 适合关注AI可信评估的研究者与政策制定者参考。

AI的迅猛发展带来巨大机遇,但也暴露了评估体系的重大缺陷。当前基准测试普遍存在数据污染、开发者选择性报告等问题,导致评估结果失真,甚至无意中偏袒特定方法。随着大量参与者涌入,评估环境如同“无序西部”,难以区分真实进步与夸大宣传,削弱科学信号并损害公众信任。类比高风险考试(如SAT、GRE)的严格监管,我们应为AI评估建立更高标准。本文主张变革现有放任式评估模式,提出一种统一、实时、质量可控的评测范式:通过密封执行、题库滚动更新与延迟透明等机制构建鲁棒框架。为此,我们推出PeerBench(https://www.peerbench.ai/),一个社区主导的评测原型系统,旨在恢复评估公信力,提供真正可信的AI进展度量。

原文摘要 · Abstract (English)

The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as current benchmarks increasingly reveal critical vulnerabilities. Issues like data contamination and selective reporting by model developers fuel hype, while inadequate data quality control can lead to biased evaluations that, even if unintentionally, may favor specific approaches. As a flood of participants enters the AI space, this "Wild West" of assessment makes distinguishing genuine progress from exaggerated claims exceptionally difficult. Such ambiguity blurs scientific signals and erodes public confidence, much as unchecked claims would destabilize financial markets reliant on credible oversight from agencies like Moody's. In high-stakes human examinations (e.g., SAT, GRE), substantial effort is devoted to ensuring fairness and credibility; why settle for less in evaluating AI, especially given its profound societal impact? This position paper argues that the current laissez-faire approach is unsustainable. We contend that true, sustainable AI advancement demands a paradigm shift: a unified, live, and quality-controlled benchmarking framework robust by construction, not by mere courtesy and goodwill. To this end, we dissect the systemic flaws undermining today's AI evaluation, distill the essential requirements for a new generation of assessments, and introduce PeerBench (with its prototype implementation at https://www.peerbench.ai/), a community-governed, proctored evaluation blueprint that embodies this paradigm through sealed execution, item banking with rolling renewal, and delayed transparency. Our goal is to pave the way for evaluations that can restore integrity and deliver genuinely trustworthy measures of AI progress.

AI评估基准测试可信评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。