统一智能体评测标准,让所有平台通用一个接口。
CUBE: A Standard for Unifying Agent Benchmarks
- 基于MCP和Gym构建分层协议,一次封装即可跨平台使用。
- 打破评测碎片化,减少重复集成工作量。
- 适合研究者与平台开发者共建标准化评测体系。
智能体评测基准的激增导致严重碎片化,威胁研究效率。每个新基准都需要大量定制集成,形成“集成税”,限制全面评估。我们提出CUBE(通用统一基准环境),一种基于MCP和Gym的通用协议标准,实现基准一次性封装、全域通用。通过将任务、基准、包和注册表等职责分离到独立API层,任何符合标准的平台均可无须定制集成地访问任意合规基准,用于评估、强化学习训练或数据生成。我们呼吁社区在2026年前共同参与标准建设,防止平台实现固化加剧碎片化。
原文摘要 · Abstract (English)
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。