为大模型智能体系统提供信任与安全评估框架
TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- 构建面向多智能体系统的信任风险安全管理体系
- 提出协作质量与工具使用效率的新评估指标
- 适合关注AI安全与合规的开发者和研究者
基于大语言模型的智能体多智能体系统正在重塑企业与社会领域的智能、自主性、协作与决策能力。本文系统分析了此类系统在信任、风险与安全管理(TRiSM)方面的挑战,提出适配智能体系统的框架,涵盖可解释性、ModelOps、安全、隐私及生命周期治理五大支柱。构建了针对智能体系统的风险分类体系,涵盖协调失败、提示攻击等独特威胁。提出两个新评估指标:组件协同得分(CSS)量化智能体间协作质量,工具利用效率(TUE)衡量工作流中工具使用效能。讨论提升可解释性、通过加密与对抗鲁棒性增强安全隐私的方法,并给出负责任发展的研究路线图,推动系统在安全性、透明性与问责性上的落地。
原文摘要 · Abstract (English)
Agentic AI systems, built upon large language models (LLMs) and deployed in multi-agent configurations, are redefining intelligence, autonomy, collaboration, and decision-making across enterprise and societal domains. This review presents a structured analysis of Trust, Risk, and Security Management (TRiSM) in the context of LLM-based Agentic Multi-Agent Systems (AMAS). We begin by examining the conceptual foundations of Agentic AI and highlight its architectural distinctions from traditional AI agents. We then adapt and extend the AI TRiSM framework for Agentic AI, structured around key pillars: \textit{ Explainability, ModelOps, Security, Privacy} and \textit{their Lifecycle Governance}, each contextualized to the challenges of AMAS. A risk taxonomy is proposed to capture the unique threats and vulnerabilities of Agentic AI, ranging from coordination failures to prompt-based adversarial manipulation. To support practical assessment in Agentic AI works, we introduce two novel metrics: the Component Synergy Score (CSS), which quantifies the quality of inter-agent collaboration, and the Tool Utilization Efficacy (TUE), which evaluates the efficiency of tool use within agent workflows. We further discuss strategies for improving explainability in Agentic AI, as well as approaches to enhancing security and privacy through encryption, adversarial robustness, and regulatory compliance. The review concludes with a research roadmap for the responsible development and deployment of Agentic AI, highlighting key directions to align emerging systems with TRiSM principles-ensuring safety, transparency, and accountability in their operation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。