早期不确定性无法预测长程智能体失败,因路径切换打断了信号与结果的关联。
Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents

- 发现路径切换是导致早期信号失效的核心机制
- 最终步信心的AUROC达0.85,早期均低于0.60
- 建议用最后一步信心决定是否重启,比中途干预更有效
长时序智能体的早期失败预测对及时干预和降低推理成本至关重要。尽管不确定性度量(如口头置信度、困惑度)被视为有前景的失败检测方法,但其在长程执行中期是否仍具判别力尚未明确。我们在深度研究任务上评估主流不确定性信号,发现口头置信度仅在轨迹完成时能可靠区分失败,平均AUROC为0.85;而在50%轨迹进度时,所有信号的平均AUROC均未超过0.60,预测能力有限。我们揭示了这一差距的根源:路径切换——智能体在执行过程中频繁放弃当前搜索方向,破坏了早期信号与最终结果之间的关联。该发现挑战了中间不确定性可指导干预的假设,并提出实用建议:在深度研究场景中,应以最终步骤的置信度判断是否重启,实验表明此法优于中途干预。
原文摘要 · Abstract (English)
Early failure prediction is important for long-horizon agents, as it enables timely intervention and can reduce inference and tool-use costs. Uncertainty quantification, such as verbal confidence and perplexity, offers a promising approach to detecting agent failures; however, it has not been explored whether these signals retain their discriminative power during the intermediate stages of long-horizon execution. We evaluate mainstream uncertainty signals on deep-research tasks and find that verbal confidence reliably distinguishes failures at trajectory completion, achieving a mean AUROC of 0.85, whereas all evaluated signals offer limited predictive value earlier in execution, with none exceeding a mean AUROC of 0.60 at 50% trajectory progress. We identify an underlying mechanism explaining this gap: path switching, where agents frequently abandon their current search direction in-trajectory, breaking the link between early signal and final outcome. These findings challenge the assumption that intermediate uncertainty can reliably guide early intervention. They also motivate a practical recommendation for agent harnesses in deep-research settings: use final-step confidence to decide whether to restart, an approach that our experiments find more effective than in-trajectory intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。