arXiv:2607.28802cs.AI2026-07

提出交互中心故障分类法,精准定位智能体失败源头并指导修复方向。

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

论文配图:Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
图 1 · 摘自论文原文
  • 按组件间交互关系划分41种故障模式,明确故障归属边与责任侧。
  • 在4个前沿模型上验证,人工与智能评判者一致性达κ=0.76,可靠性高。
  • 适用于编码助手、长期任务助理等多类智能体系统,修复可操作性强。

现有评估常将智能体失败简化为系统级结果,掩盖了故障源头及对应修复方式。同一可见失败可能源于模型后训练、工具集成、环境设计或评测标准问题,导致修复归因困难。由于智能体行为源自模型、工具、记忆、用户、环境等组件间的复杂交互,仅依赖结果标签难以改进。多数故障分类法因局限于特定基准且缺乏统一结构而无效。本文提出一种以交互为中心的分类体系,将失败定位到具体交互环节,并标识责任组件。该体系将41种故障模式映射至组件间连接边及故障侧,实现可操作性:模型侧故障指向后训练优化,工具侧故障提示结构与集成修复,环境或评分器故障则暴露需重构的评估条件。该框架适用于各类智能体架构,从代码生成到长周期个人助手及多智能体系统。通过公共基准、模型系统卡、报告和日志轨迹中的实例进行实证,使用独立推理智能体作为评判者评估可复现性。在四个前沿模型上,最强评判者与人工标签的克朗巴赫κ值达0.76,表明分类体系捕捉的是共享结构而非标注偏好。

原文摘要 · Abstract (English)

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model post-training, harness engineering, environment redesign, or benchmark repair depending on its source. Because agent behavior emerges from interactions among models, harnesses, users, tools, memory, and environments, outcome-level labels are often insufficient for improvement. Most failure taxonomies do little to resolve this problem because they are benchmark-specific and lack a shared structure. We introduce an interaction-centric taxonomy that localizes failures to the interactions in which they originate and identifies the responsible component. It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. This makes the taxonomy actionable: model-side failures identify targets for post-training, harness-side failures point to scaffolding and tool-integration fixes, and environment or grader failures reveal evaluation conditions requiring redesign. The schema applies across agent architectures, from coding assistants to long-horizon personal assistants and multi-agent systems. We ground the taxonomy in worked examples from public benchmarks, model system cards, published reports, and logged agent trajectories, and evaluate its reproducibility using independent reasoning agents as judges. Across four frontier models, the strongest judge reaches Cohen's $κ=0.76$ against human category labels, suggesting that the categories capture shared structure rather than annotator-specific preferences.

智能体故障诊断交互分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。