arXiv:2605.14164cs.AI2026-05

AI模型发布常选特定基准测试,背后隐藏着竞争叙事而非科学标准。

Unsteady Metrics and Benchmarking Cultures of AI Model Builders

论文配图:Unsteady Metrics and Benchmarking Cultures of AI Model Builders
图 1 · 摘自论文原文
  • 分析139次模型发布中的231个基准测试,发现多数仅被单一厂商使用
  • 63.2%的基准仅由一个厂商使用,38.5%只出现一次,跨模型可比性极低
  • 基准被用于塑造市场形象,而非客观评估,常以泛化知识为名测数学等具体能力

当前基础与生成式AI模型的能力评价已从学术论文转向企业发布会和博客,模型开发者选择性展示特定基准测试结果来定义技术前沿。我们构建并开源了Benchmarking-Cultures-25数据集,包含2025年11家主要AI厂商发布的139次模型更新中提及的231个基准测试,并提供交互式探索工具。分析显示评估体系高度碎片化:63.2%的基准仅被一个厂商使用,38.5%仅出现在一次发布中。仅有少数基准如GPQA Diamond、LiveCodeBench、AIME 2025获得广泛采用。不同厂商对同一基准赋予不同能力标签,其解释依赖于自身叙事。我们提出统一分类体系,将多元术语映射至基准作者声称测量的核心信号。其中“通用知识应用”是最常见但最模糊的类别;定性分析发现,多数基准忽视建构效度,转而将结果作为通向AGI进展的指标。尽管宣称衡量知识或推理,实际多聚焦于STEM领域(尤其是数学)。我们主张,被突出的基准更多是灵活的叙事工具,服务于市场定位,而非标准化测评手段。

原文摘要 · Abstract (English)

The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, where model builders highlight results on selected benchmarks. These artifacts now largely define the state of the art for researchers and the public. Despite their prominence, which benchmarks model builders choose to highlight, and what they communicate through this selection, is underexamined. To investigate, we introduce and open-source Benchmarking-Cultures-25, a dataset of 231 benchmarks highlighted across 139 model releases in 2025 from 11 major AI builders, alongside an interactive tool to explore the data. Our analysis reveals a fragmented evaluation landscape with limited cross-model comparability: 63.2% of highlighted benchmarks are used by a single builder, and 38.5% appear in just one release. Few achieve widespread use (e.g., GPQA Diamond, LiveCodeBench, AIME 2025). Moreover, benchmarks are attributed different competencies by different builders, depending on their narrative. To disentangle these conflicting presentations, we develop a unified taxonomy mapping diverging terminology to a shared framework of measured signals based on what benchmark authors claim to measure. "General knowledge application" is the second most popular, yet vaguely defined, category. Qualitative analysis shows many such benchmarks deemphasize construct validity, instead framing results as indicators of progress toward AGI. Their authors claim to measure knowledge or reasoning broadly, yet mostly evaluate STEM subjects (especially math). We argue that highlighted benchmarks function less as standardized measurement tools and more as flexible narrative devices prioritizing market positioning over scientific evaluation. Data: https://hf.co/datasets/matybohacek/benchmarking-cultures-25; tool: https://bench-cultures.net.

AI评估基准测试模型发布叙事机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。