提出长轨迹故障诊断新基准,精准定位失败根源与责任角色。
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

- 基于段落摘要检索错误步骤,无需训练直接追溯源头
- 最长轨迹145步,最先进方法根因定位准确率仅13.2%
- 新方法在责任角色识别上达51.1%,根因步骤准确率24.1%
当长周期智能体执行失败时,结果层面的评估只能显示失败结果,无法定位关键错误发生的位置。开发者需人工检查完整执行过程以确定责任角色并定位最早导致失败的根因步骤。现有故障归因基准多聚焦于短轨迹,对包含数百步记录的长轨迹诊断研究不足。我们提出LongRCA Bench,包含五个领域共1,140条无注入错误的失败轨迹,提供独立评分的人工标注的责任角色和最早根因步骤标签。中位轨迹长度为145步,最强基线模型在根因步骤精确匹配上仅达13.2%准确率。我们进一步提出无需训练的根因轨迹归因方法(RCTA),通过段落摘要检索候选错误步骤,并回溯至早期交接指令。在相同骨干模型、基准实例和评分协议下,RCTA实现51.1%的责任角色准确率和24.1%的根因步骤精确匹配率。结果表明,在长轨迹故障诊断中,应将责任角色归因与根因步骤定位视为独立目标进行评估。
原文摘要 · Abstract (English)
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。