构建可信赖的智能体评估体系,解决现有评测不透明、安全遗漏等问题。
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

- 通过执行轨迹、审计日志、环境快照三通道记录,实现过程可追溯的评分
- 发现44%安全漏洞和13%鲁棒性问题在传统评测中被忽略
- 揭示模型能力多维性,适合关注可靠部署的开发者与研究者
大型语言模型正越来越多地作为自主智能体,在真实软件环境中执行多步骤任务。然而,现有智能体评测存在轨迹不可见评分、安全与鲁棒性评估不明确、模态与交互范式覆盖窄等局限。我们提出Claw-Eval,一个端到端评估套件,包含300个经人工验证的任务,覆盖9个类别,分为三大组:通用服务编排、多模态感知与交互、多轮专业对话。为实现轨迹感知评分,每轮运行通过执行轨迹、审计日志、环境快照三个独立证据通道记录,生成2,159个细粒度评分项。评分协议评估完成度、安全性与鲁棒性,使用平均分、Pass@k及Pass^k(三次试验)区分真实能力与偶然成功。对14个前沿模型的实验表明:(1) 轨迹不可见评测系统性不可靠,漏检44%安全违规和13%鲁棒性失败;(2) 能力不等于一致性,错误注入下Pass@3稳定,而Pass^3下降最高达24个百分点;(3) 智能体能力高度多维,不同任务组与指标下模型排名差异显著,证明评估覆盖多样性至关重要。Claw-Eval指明了开发不仅强大且可信赖智能体的方向。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safety and robustness evaluation, and narrow coverage of modalities and interaction paradigms. We introduce Claw-Eval, an end-to-end evaluation suite addressing these gaps with 300 human-verified tasks spanning 9 categories across three groups: general service orchestration, multimodal perception and interaction, and multi-turn professional dialogue. To enable trajectory-aware grading, each run is recorded through three independent evidence channels: execution traces, audit logs, and environment snapshots, yielding 2,159 fine-grained rubric items. The scoring protocol evaluates Completion, Safety, and Robustness, with Average Score, Pass@k, and Pass^k across three trials to distinguish genuine capability from lucky outcomes. Experiments on 14 frontier models show that: (1) Trajectory-opaque evaluation is systematically unreliable, missing 44% of safety violations and 13% of robustness failures detected by our framework. (2) Capability does not imply consistency, with Pass@3 remaining stable under error injection while Pass^3 dropping by up to 24 percentage points. (3) Agent capability is strongly multi-dimensional, with model rankings varying across task groups and metrics, indicating that our heterogeneous evaluation coverage is essential. Claw-Eval highlights directions for developing agents that are not only capable but reliably deployable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。