提出用预测有效性评估大模型智能体,解决排行榜在实际部署中失效的问题。
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
- 以预测有效性替代平均分排序,衡量模型在新场景下的表现稳定性。
- 14项并行实验揭示现有评测忽略的关键维度,如多模态、推理模式等。
- 提出可验证的三重外部测试标准,适合工业级智能体研发与评估者参考。
智能体评测基准发展迅速,但单一基准最多覆盖五维部署场景。本文整合迄今最大规模的基于MCP的工业级智能体评测:14项并行实现研究,涵盖新资产类别(包括多模态视觉扩展)、不同编排方式、检索策略、推理模式、基础设施优化及评估方法探针。结合此前7个智能体基准,我们指出聚合得分排行榜系统性地低估了实际部署需求;基于样本内排名的得分无法外推至分布外场景,公开竞赛转为私有后的实证数据直接支持这一排名不稳定性。我们主张采用预测有效性——即样本内与样本外排名的相关性——作为排序依据,并报告一个十二级测量体系,揭示HELM及其后续智能体时代基准所忽视的部署相关维度。该观点通过三个可验证的分布外标准予以操作化,明确阈值;现有证据部分支持但尚不足以确认。最后,提出预注册试点设计及下一代智能体评测的领域愿景。
原文摘要 · Abstract (English)
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。