arXiv:2608.16411cs.SEcs.AI2026-08

通过分析智能体行为轨迹,实现安全可靠的部署

Towards Risk-free AI Agent Deployment

论文配图:Towards Risk-free AI Agent Deployment
图 1 · 摘自论文原文
  • 以智能体的推理与操作轨迹为依据进行风险检测
  • 提出覆盖全生命周期的部署就绪检查清单
  • 适合关注AI代理安全与可维护性的研发团队

基于大模型的智能体正快速从研究原型进入企业核心流程,但其部署存在安全、合规与功能风险。本文主张风险防控应基于智能体的完整轨迹——包括推理步骤、工具调用与环境观测序列。这些轨迹可记录且多数故障仅在轨迹中显现。为实现可持续部署,我们倡导将智能体测试与调试作为系统性研究方向。文章首先分析测试挑战:验证难题、非确定性、轨迹校验及缺乏充分性度量。随后探讨调试方法,涵盖自动故障归因、修复与自我演进。最后提炼出贯穿部署全周期的就绪检查清单,并指出三大开放问题:形式化充分性度量、长程轨迹根因定位、自进化智能体的可靠性,需社区共同解决以实现可信部署。

原文摘要 · Abstract (English)

LLM-based agents are rapidly moving from research prototypes into the core business processes of organizations, but these agents pose deployment risks to security, compliance, and functionality. In this article, we argue that risk-free deployment must be grounded in the agent's trajectory: the recorded sequence of reasoning steps, tool invocations, and environmental observations. Trajectories are available for any agent, and many failures are visible only in the trajectory. To make agents deployable and sustainable, we advocate agent testing and debugging as a systematic research direction for detecting and mitigating these risks. This article begins with the challenges of testing agents, including the oracle problem, non-determinism, trajectory validation, and the absence of adequacy metrics. We then turn to debugging agents, from automated failure attribution to repair and self-evolution. We distill these directions into a practical deployment-readiness checklist covering the full deployment lifecycle. Finally, we identify open problems, i.e., formal adequacy metrics, root-cause attribution over long-horizon trajectories, and the reliability of self-evolving agents, that the community must address to enable trustworthy agent deployment.

AI代理风险控制部署安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。