arXiv:2602.19008cs.CLcs.LG2026-02被引 3

语言智能体失败主因是路径漂移,而非能力不足。

Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks

  • 用工具使用任务的规范路径作为基准,衡量轨迹偏离程度。
  • 偏离规范路径的执行导致成功率下降,且偏差会自我强化。
  • 通过中途监测并重启低合规性任务,可提升8.8%成功率。

为何语言智能体在具备解决能力的任务上仍会失败?我们提出,许多失败源于随机漂移偏离任务潜在解结构的可靠性问题,而非能力不足。每个定义明确的工具使用任务都存在一条规范解路径(即成功运行中共享的一组工具调用序列),而智能体成败关键在于轨迹是否保持在此路径的运作区间内。通过一个自然实验,固定模型能力和任务难度,分析Toolathlon基准数据:22个前沿模型在3次独立运行中尝试108个真实世界工具使用任务,共产生515个模型×任务单元,同一模型在不同运行中有的成功、有的失败,仅由大模型采样随机性造成。在这些单元中,成功运行显著更贴近规范解路径(Jaccard值高0.060,p<0.0001,n=488,95%置信区间[+0.043, +0.077]),该结果经六项稳健性检验验证。关键发现:该因果机制是渐进且自增强的——前50%轨迹中偏差无统计差异,排除早期分支偏见;每次偏离规范路径的工具调用使下一次偏离概率上升22.7个百分点(β̂=+0.227,p<0.0001),超过基线率一倍以上。这表明仅靠提升能力无法改善可靠性,但可采取行动:在中段轨迹引入简单监控,对合规性最低的三分之一运行进行重启,可使干预组成功率提升8.8个百分点。

原文摘要 · Abstract (English)

Why do language agents fail on tasks they are capable of solving? We argue that many such failures are reliability failures caused by stochastic drift from a task's latent solution structure, not capability failures. Every well-defined tool-use task imposes a canonical solution path (i.e., a convergent set of tool invocations shared across successful runs) and agent success depends critically on whether a trajectory stays within this path's operating envelope. We establish this causally using a natural experiment that holds model capability and task difficulty fixed by construction. We analyze trajectories from the Toolathlon benchmark: 22 frontier models each attempt 108 real-world tool-use tasks across 3 independent runs, yielding 515 model$\times$task units where the same model succeeds on some runs and fails on others due to LLM sampling stochasticity alone. Within these units, successful runs adhere significantly more closely to the canonical solution path than failed runs ($+$0.060 Jaccard, $p<0.0001$, $n=488$ units, 95% CI [+0.043, +0.077]). This result survives six robustness checks including cross-model-family leave-one-out validation. Critically, the causal mechanism is gradual and self-reinforcing: the adherence gap is statistically indistinguishable from zero through the first 50% of the trajectory, ruling out early-branching selection bias, and each off-canonical tool call raises the probability that the next call is also off-canonical by 22.7 percentage points ($\hatβ=+0.227$, $p<0.0001$), more than doubling the baseline rate. These findings imply that agent reliability cannot be improved by capability scaling alone, but offer a highly actionable intervention: a simple monitor that restarts the bottom tercile of runs based on mid-trajectory canonical adherence lifts success rates by $+$8.8 percentage points among intervened runs.

智能体可靠性路径偏差工具使用大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。