arXiv:2606.13608cs.AIcs.LG2026-06被引 1

用智能体评估智能体,让评测更开放、标准且可复现。

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

论文配图:AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
图 1 · 摘自论文原文
  • 评测由裁判智能体执行,通过统一协议交互,解耦评估逻辑与实现。
  • 五个月竞赛中300+裁判智能体验证了跨异构场景的兼容性。
  • 适合追求公平、可复现评测的研究者和开发者使用。

智能体系统快速发展,但评估仍碎片化。现有基准多依赖固定的大模型驱动框架,需大量集成,产生测试-生产差异,限制不同智能体设计间的公平比较。根本问题在于缺乏开放、通用的评估接口。本文倡导智能体化智能体评估(AAA),由裁判智能体通过标准化协议A2A(任务管理)和MCP(工具访问)进行评测。传统基准需两个独立接口,而AAA仅需一个,形成统一框架,分离评估逻辑与实现,支持可复现、互操作、多智能体评测。我们进一步提出AgentBeats作为实现:识别五种实用运行模式,适配开放性、隐私和可复现性的现实约束。通过两项研究验证:为期五个月的公开竞赛汇聚12类共298个裁判智能体和467个被测智能体;编码智能体案例研究证实,该方法保持与公共记录的一致性,并揭示此前缺失的直接对比结果,带来关于智能体设计的新洞见。结合大规模实地研究与受控案例,证明AAA在异构场景下具备覆盖性、实用性与保真度。整体上,AAA与AgentBeats为开放、标准、可复现的智能体评估提供了清晰路径。

原文摘要 · Abstract (English)

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.

智能体评测标准化可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。