arXiv:2602.11409cs.AI2026-02被引 8

提出轨迹级不确定性度量TRACER,精准识别智能体交互中的关键失败时刻。

TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning

  • 融合内容意外性与情境感知信号,捕捉对话中潜在风险。
  • 在τ²-bench上使任务失败预测的AUROC提升37.1%,检测更早更准。
  • 适合需要高可靠性人机协作的复杂工具使用场景研究者使用。

在真实世界多轮人机协同使用工具的交互中,估计智能体的不确定性极具挑战,因为失败常由稀疏的关键事件(如循环、不连贯的工具调用或用户-智能体协调失误)引发,即使局部生成表现自信。现有不确定性指标集中于单次文本生成,无法捕捉此类轨迹级异常信号。本文提出TRACER,一种面向双控(用户-智能体-工具)交互的轨迹级不确定性度量。TRACER结合内容感知的意外性、情境感知信号、语义与词汇重复性,以及基于工具的连贯性缺口,并通过尾部聚焦的风险函数与最大复合步骤风险聚合机制,突出显示决定性异常。在τ²-bench上评估其对任务失败预测与选择性任务执行的能力,结果表明,TRACER相较基线方法,最高提升AUROC达37.1%,提升AUARC达55%,显著增强了复杂对话式工具使用场景下的不确定性检测能力。代码与基准已开源。

原文摘要 · Abstract (English)

Estimating uncertainty for AI agents in real-world multi-turn tool-using interaction with humans is difficult because failures are often triggered by sparse critical episodes (e.g., looping, incoherent tool use, or user-agent miscoordination) even when local generation appears confident. Existing uncertainty proxies focus on single-shot text generation and therefore miss these trajectory-level breakdown signals. We introduce TRACER, a trajectory-level uncertainty metric for dual-control Tool-Agent-User interaction. TRACER combines content-aware surprisal with situational-awareness signals, semantic and lexical repetition, and tool-grounded coherence gaps, and aggregates them using a tail-focused risk functional with a MAX-composite step risk to surface decisive anomalies. We evaluate TRACER on $τ^2$-bench by predicting task failure and selective task execution. To this end, TRACER improves AUROC by up to 37.1% and AUARC by up to 55% over baselines, enabling earlier and more accurate detection of uncertainty in complex conversational tool-use settings. Our code and benchmark are available at https://github.com/sinatayebati/agent-tracer.

不确定性估计智能体交互轨迹分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。