为复杂企业级智能体系统构建动态安全框架,识别并应对新型风险。
A Safety and Security Framework for Real-World Agentic Systems
- 通过用户安全视角重构智能体风险,融合传统安全与新兴风险。
- 在真实场景中发现10,000+次攻击与防御交互,验证框架有效性。
- 支持企业级部署,适合关注智能体安全的研究者与开发者。
本文提出一种动态可操作的安全框架,用于保障企业级智能体系统的安全性。我们认为,安全与风险并非单一模型的固有属性,而是由模型、编排器、工具和数据在运行环境中的动态交互所涌现的特性。通过用户安全视角,我们发现传统上分离的安全与风险在智能体系统中紧密关联。基于此,构建了一个统一的智能体风险分类体系,涵盖工具滥用、连锁动作链、意外控制放大等新型风险。核心是采用辅助AI模型与智能体,在人类监督下实现上下文感知的风险发现、评估与缓解。针对最难点——风险探测,提出沙箱化、AI驱动的红队测试方法。通过案例研究验证了该框架在NVIDIA旗舰智能体研究助手AI-Q中的应用,完成端到端的安全评估,发现了多个新风险并实现上下文化解。同时发布包含超10,000条真实攻击与防御行为轨迹的数据集,推动智能体安全研究发展。
原文摘要 · Abstract (English)
This paper introduces a dynamic and actionable framework for securing agentic AI systems in enterprise deployment. We contend that safety and security are not merely fixed attributes of individual models but also emergent properties arising from the dynamic interactions among models, orchestrators, tools, and data within their operating environments. We propose a new way of identification of novel agentic risks through the lens of user safety. Although, for traditional LLMs and agentic models in isolation, safety and security has a clear separation, through the lens of safety in agentic systems, they appear to be connected. Building on this foundation, we define an operational agentic risk taxonomy that unifies traditional safety and security concerns with novel, uniquely agentic risks, including tool misuse, cascading action chains, and unintended control amplification among others. At the core of our approach is a dynamic agentic safety and security framework that operationalizes contextual agentic risk management by using auxiliary AI models and agents, with human oversight, to assist in contextual risk discovery, evaluation, and mitigation. We further address one of the most challenging aspects of safety and security of agentic systems: risk discovery through sandboxed, AI-driven red teaming. We demonstrate the framework effectiveness through a detailed case study of NVIDIA flagship agentic research assistant, AI-Q Research Assistant, showcasing practical, end-to-end safety and security evaluations in complex, enterprise-grade agentic workflows. This risk discovery phase finds novel agentic risks that are then contextually mitigated. We also release the dataset from our case study, containing traces of over 10,000 realistic attack and defense executions of the agentic workflow to help advance research in agentic safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。