揭示智能体任务成功背后的运行异常,提供可量化的审计数据集。
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

- 构建真实智能体执行轨迹的异常标注数据集,关联任务结果与过程证据。
- 31,135次任务通过中仍有2,904次被识别为过程异常,说明成功评价有盲区。
- 提供结构化异常标签与检测模型,适合研究智能体可靠性与安全监控。
任务成功可能掩盖真实智能体执行中的过程异常。智能体虽通过最终任务验证,但仍可能存在未解决的歧义、不安全外部写入、忽略错误、弱根基承诺或能力边界过度承诺等问题。我们将其称为‘结果-过程差距’,并提出 OpenClawBench,一个大规模数据集,用于测量和监督智能体执行过程中的过程侧异常。该数据集基于 6 个源模型通过 BFCL 驱动的 OpenClaw 会话生成,包含 31,264 条标注轨迹,实现任务验证结果与结构化过程证据对齐。FullTax 将对齐轨迹转化为结构化异常监督信息:二值标签、支持证据、异常发生/持续时间定位、严重性、可恢复性及五类异常分类。利用 OpenClawBench,我们使结果-过程差距可量化:在 31,135 次通过任务验证的执行中,仍有 2,904 次被判定为过程异常。这表明仅以成功为导向的评估会遗漏一类真实存在的过程失败。基于高置信度 FullTax 监督池微调的 LoRA-Gemma 3 12B 检测器,在干净标签测试集上达到二值 F1=0.729。OpenClawBench 将真实智能体执行日志转化为可审计、可复用的监督信号,支持运行时智能体可靠性研究、诊断与监控。
原文摘要 · Abstract (English)
Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external writes, ignored errors, weakly grounded commitments, or capability-boundary overcommitment. We study this mismatch as the Outcome-Process Gap and introduce OpenClawBench, a large-scale dataset for measuring and supervising process-side anomalies in real agent execution processes. OpenClawBench is built from BFCL-driven OpenClaw sessions produced by 6 source models and contains 31,264 annotated trajectories. It aligns task-oracle outcomes with structured process evidence. FullTax converts the aligned trajectories into structured anomaly supervision: binary labels, supporting evidence, onset/span localization, severity, recoverability, and a 5-class anomaly taxonomy. Using OpenClawBench, we make the Outcome-Process Gap measurable. Among 31,135 oracle-passing executions, 2,904 are still labeled process-anomalous under FullTax. These results show that success-only evaluation misses a concrete class of process-side failures in real agent executions. A LoRA-fine-tuned Gemma 3 12B detector trained on the high-confidence FullTax supervised pool reaches binary F1=0.729 on the cleaner-labels held-out test split. Together, OpenClawBench turns real agent execution logs into auditable and reusable supervision for studying, diagnosing, and operationally monitoring runtime agent reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。