arXiv:2608.04018cs.CYcs.AI2026-08

通过追踪执行轨迹发现并阻断智能体系统风险,提升组织级AI安全性。

Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming

论文配图:Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
图 1 · 摘自论文原文
  • 基于执行轨迹识别多步推理中的潜在漏洞
  • 在四个任务套件上使攻击成功率降至接近零
  • 适合关注AI安全与合规的组织管理者

智能体正越来越多地嵌入组织工作流中,通过交互外部信息源和调用数字工具完成操作任务。随着系统应用深入,来自恶意或不可信外部信息的风险日益突出,可能引导智能体产生非预期行为。现有红队测试方法多依赖固定攻击模板或最终攻击结果,难以揭示攻击在多步推理与工具使用过程中的演变路径。本文认为,智能体执行风险应从轨迹层面理解。基于此,提出TrajRed框架,利用执行轨迹挖掘智能体系统的漏洞;进一步开发了运行时治理层TrajGuard,基于红队发现的高风险轨迹实时监控并干预任务流程。在AgentDojo平台上对四个组织任务套件的实验表明,TrajRed识别出的漏洞显著强于固定模板与自动红队基线。依托这些发现,TrajGuard将所有评估攻击方法的成功率降至接近零,同时保持正常任务功能。结果表明,执行轨迹为红队测试与风险控制提供了切实可行的基础,凸显了在组织级AI部署中治理执行过程的重要性。

原文摘要 · Abstract (English)

AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is identifying and mitigating risks arising from malicious or untrusted external information that can steer agents toward unintended actions. Existing red-teaming approaches largely rely on fixed attack templates or final attack outcomes, providing limited visibility into how attacks unfold through multi-step reasoning and tool use. We argue that agent execution risk should be understood as a trajectory-level phenomenon. Building on this perspective, we propose TrajRed, a trajectory-guided red-teaming framework that uses execution trajectories to uncover vulnerabilities in agentic AI systems. We further develop TrajGuard, a runtime governance layer that uses high-risk trajectories discovered during red teaming to monitor and intervene in ongoing workflows. Experiments on AgentDojo across four organizational task suites show that TrajRed identifies substantially stronger vulnerabilities than fixed-template and automatic red-team baselines. Building on these vulnerability findings, TrajGuard reduces attack success across all evaluated attack methods to near zero while preserving benign task utility. Together, the results demonstrate that execution trajectories provide a practical foundation for both red teaming and risk control in agentic AI systems. This work highlights the importance of governing agent execution in organizational AI deployments.

AI安全智能体红队测试风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。