arXiv:2608.22676cs.AI2026-08

通过多信号融合,精准识别工具返回异常并判断其影响。

Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

  • 在工具返回时融合概率与动作分布信号,定位失败类型。
  • 似然信号有效捕捉不完整和部分不一致结果,准确率超90%。
  • 适合开发高可靠智能体的工程师与研究者参考。

使用外部工具的智能体在执行多步任务时,工具返回可能以不同方式失败,需不同恢复策略。现有鲁棒性研究多依赖不确定性指标检测代理不可靠性,但无法直接识别失败类型或应对方式。本文在工具返回进入代理上下文的瞬间分析失败,结合两种互补信号:一是返回内容在工具模式下的似然与完整轨迹前缀下的似然对比;二是代理对合法下一步动作的概率分布。我们在零售客服基准中注入不完整与不一致的返回进行评估,结果显示,似然信号能清晰捕捉不完整返回及部分直接不一致,而动作信号揭示失败对后续决策强度的影响。某些在似然信号下较弱的失败仍可引导代理转向状态变更动作。结果表明,工具失败可在返回边界被识别,但可靠诊断需多信号融合。

原文摘要 · Abstract (English)

Tool-using agents increasingly rely on external tools to complete multi-step tasks, but tool returns can fail in different ways and require different recovery actions. Existing robustness studies often use uncertainty-based measures to detect when an agent becomes unreliable. These measures can reveal that something has gone wrong, but they do not directly identify the type of tool failure or the appropriate response. We address this limitation by analyzing tool failures at the moment a return enters the agent context. Our approach combines two complementary signals. The first compares the likelihood of the returned content under the tool schema and under the full trajectory prefix. The second measures the agent's probability distribution over its legal next actions. We evaluate the approach by injecting incomplete and inconsistent returns into a retail customer-service benchmark. The results show that likelihood-based signals clearly capture incomplete returns and some direct inconsistencies, while action-based signals reveal how strongly a failure changes the next decision. Some failures that are weak under likelihood signals can still redirect the agent toward state-changing actions. These findings show that tool failures can be recognized at the return boundary, but reliable diagnosis requires combining multiple signals.

智能体工具使用鲁棒性故障检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。