诊断长时序安全大模型代理失败原因,发现早期状态遗漏是主因。
Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

- 通过检查点分离任务失败阶段,定位上游瓶颈
- 引导式提示使状态观测率从65.5%提升至95.4%
- 揭示不同模型版本失败模式差异,适合安全评估者参考
长时序安全大模型代理需在多步依赖交互中保持信息与决策连贯性,后继动作常依赖早期发现的服务、状态或权限。这使得最终任务成功难以解释:代理可能在能力暴露前已失败。本文提出诊断方法,在安全任务中插入检查点,区分能力暴露前后的失败,并通过可控干预测试潜在上游瓶颈。在四个任务族中验证:延迟信息重用、状态重用、策略失败恢复、不确定结果后的决策。状态重用任务中,检查点分析显示,众多Gemini 2.5 Flash的失败发生在其观察目标状态之前。在预设的92种子实验中,针对性协议消歧引导使状态观测率从对照组的65.5%提升至95.4%。相同设计应用于Gemini 3.7 Flash却产生相反效果,且状态观测不再可靠预测任务完成。结果表明,失败主因随模型版本变化,强调应诊断失败位置与原因,而非仅依赖总体成功率。
原文摘要 · Abstract (English)
Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success difficult to interpret: an agent may fail before it ever reaches the point where the capability of interest can be exercised. We present a diagnostic methodology that instruments security tasks with checkpoints, separates failures before and after capability exposure, and uses controlled interventions to test suspected upstream bottlenecks. We evaluate the methodology across four task families involving delayed reuse of discovered information, reuse of observed state, recovery from failed strategies, and decision making after uncertain outcomes. On observed state reuse, checkpoint analysis shows that many Gemini 2.5 Flash failures occur before the model observes the state it is later expected to reuse. In a pre-specified 92-seed study, targeted protocol-disambiguation guidance increases state observation from 65.5\% under a matched non-guidance control message to 95.4\%. Repeating the same design with Gemini 3.7 Flash produces the opposite effect, while state observation no longer reliably predicts task completion. These results show that the dominant source of failure can shift across model generations, motivating evaluation that diagnoses where and why long-horizon security agents fail rather than relying only on aggregate task success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。