构建统一评估框架,让大模型智能体评测更高效透明。
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

- 分拆评测流程为基准、工具、环境三组件,灵活可配置。
- 支持20+基准测试,覆盖五大能力维度,提升可复现性。
- 内置故障容错与轨迹分析,适合研究者快速诊断问题。
随着大语言模型向自主智能体演进,统一的评估基础设施日益关键。然而当前评估流程高度碎片化且耦合紧密,阻碍了结果复现并导致重复工程。为此,我们提出 AgentCompass,一个开源、轻量且可扩展的 LLM 智能体评估框架。该框架将评估过程划分为独立的三个组件:基准(Benchmark)、工具(Harness)和环境(Environment),实现无需重写复杂执行逻辑的灵活配置。同时,其具备容错的异步运行时与全面的轨迹分析工具,可透明诊断奖励劫持等细微失败模式。原生支持超过20个基准测试,覆盖五个能力维度,为社区提供可扩展、可复现的智能体研究基础设施。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。