arXiv:2605.23989cs.AIcs.CL2026-05综述被引 4

系统梳理智能体AI在安全、鲁棒性、隐私与系统安全方面的风险与应对策略。

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

  • 从任务规划到工具调用全程分析信任风险点,提出阶段化防御方法
  • 构建统一评估体系,涵盖结果与过程信号,支持部署决策
  • 适合高风险场景下研发可信智能体系统的研究人员与工程师

自主执行复杂任务的智能体AI(由大语言模型增强规划、工具使用、记忆和长周期交互能力构成)虽具潜力,但其多步行为轨迹引入新故障模式,威胁可信性。本综述聚焦高风险应用中至关重要的两大维度:安全与鲁棒性,以及隐私与系统安全。针对每个维度,厘清核心概念,识别代理工作流中的风险环节,并总结分阶段缓解策略。其他可信性方面(价值对齐、透明度、公平性、可问责性)作为背景讨论。为支持一致评估与部署决策,我们整合出统一的指标与基准集,强调结果与过程信号(如约束违规、轨迹完整性、对抗成功率),并提供场景到指标的指导以用于发布门控。最后指出开放挑战:自演化智能体、运行时监控与验证、隐私保护个性化及信任-效用权衡,并通过开源智能体系统的实际安全失败案例进行说明。目标是为高风险环境中构建可信智能体系统的研究者与实践者提供实用参考。

原文摘要 · Abstract (English)

Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their multi-step trajectories introduce new failure modes that challenge trustworthiness. This survey provides a focused examination of trustworthy agentic AI through two core dimensions that are critical for high-risk deployments: Safety and Robustness, and Privacy and System Security. For each dimension, we clarify key concepts, identify where risks emerge along the agent workflow, and summarize stage-targeted mitigation strategies. Other trustworthiness aspects (value alignment, transparency, fairness, and accountability) are discussed as relevant context rather than parallel chapters. To support consistent comparison and deployment decisions, we consolidate evaluation into a unified metrics-and-benchmarks hub, emphasizing both outcome and process signals (e.g., constraint violations, trace completeness, and adversarial success rates) and offering scenario-to-metric guidance for release gating. We conclude by outlining open challenges such as self-evolving agents, runtime monitoring and verification, privacy-preserving personalization, and the trust-utility trade-off, and present a case study of real-world security failures in open-source agentic systems. Our goal is to serve as a practical reference for researchers and practitioners building trustworthy agentic systems in high-stakes environments.

智能体AI安全评估可信计算隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。