arXiv:2607.25891cs.AIcs.DB2026-07

构建首个跨基准的高分辨率智能体评估语料库,支持大规模性能分析。

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

  • 统一30个基准、745个智能体的评估数据,形成标准化格式
  • 发现编程能力提升最快,企业工作流仍最困难,函数调用已趋饱和
  • 可按领域、任务类型等维度估算能力得分,适合评估设计者使用

现有AI智能体评估因任务、验证器和评分规则分散而难以全面比较。我们提出MESSIER,一个包含957,611条记录的统一语料库,覆盖30个基准、745个智能体、11,891个任务和74,263个验证器。该语料库整合公开结果并新增6个代表性专业与科学基准的新运行数据,将异构组件标准化。分析显示,前沿进展在不同基准组间不均衡:函数调用评估基本饱和,编程能力提升最快,企业工作流仍最困难。反事实重评分表明,多验证器任务中严格全通过评分会改变智能体排名。基于语料库计算的能力得分与Epoch评估能力指数相关性达Spearman ρ = 0.84,且可针对领域、职业、动作空间或验证器类型子集估算。总体而言,MESSIER是可复用的智能体性能研究资源,也为设计更优评估提供基础。

原文摘要 · Abstract (English)

Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations.

智能体评估跨基准数据语料库能力量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。