提出可诊断风险根源的AI代理安全框架,提升透明度与可控性。
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- 构建三维风险分类体系,从来源、故障模式和后果三方面系统化归类风险。
- 在复杂交互场景中实现顶尖的安全调控效果,能精准识别异常行为根源。
- 支持多尺寸模型部署,适合需要高透明度的AI代理开发与安全验证者。
AI代理的兴起带来了由自主工具使用和环境交互引发的复杂安全与隐私挑战。现有防护模型缺乏对代理风险的感知能力以及风险诊断的透明性。为此,我们首先提出一个统一的三维风险分类体系,从风险来源(where)、故障模式(how)和后果(what)三个维度进行正交划分。基于此结构化层级分类,我们构建了细粒度的代理安全基准测试(ATBench),并设计了代理安全与防护诊断框架AgentDoG。AgentDoG可对代理行为轨迹进行细粒度上下文监控,更重要的是能够诊断不安全行为及看似合理但不合理行为的根本原因,提供可追溯性与透明性,超越传统二值标签,助力有效对齐。AgentDoG提供三种参数规模(4B、7B、8B)的变体,适配Qwen与Llama模型系列。大量实验表明,AgentDoG在多样且复杂的交互场景中实现了当前最优的代理安全管控性能。所有模型与数据集均已开源。
原文摘要 · Abstract (English)
The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current guardrail models lack agentic risk awareness and transparency in risk diagnosis. To introduce an agentic guardrail that covers complex and numerous risky behaviors, we first propose a unified three-dimensional taxonomy that orthogonally categorizes agentic risks by their source (where), failure mode (how), and consequence (what). Guided by this structured and hierarchical taxonomy, we introduce a new fine-grained agentic safety benchmark (ATBench) and a Diagnostic Guardrail framework for agent safety and security (AgentDoG). AgentDoG provides fine-grained and contextual monitoring across agent trajectories. More Crucially, AgentDoG can diagnose the root causes of unsafe actions and seemingly safe but unreasonable actions, offering provenance and transparency beyond binary labels to facilitate effective agent alignment. AgentDoG variants are available in three sizes (4B, 7B, and 8B parameters) across Qwen and Llama model families. Extensive experimental results demonstrate that AgentDoG achieves state-of-the-art performance in agentic safety moderation in diverse and complex interactive scenarios. All models and datasets are openly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。