用少样本训练轻量级智能体安全框架,实时防护复杂交互风险。
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

- 基于影响函数净化构建分类引导数据引擎,仅用1000样本训练小参数模型。
- 在多环境交互中性能媲美闭源大模型(如GPT-5.4),部署开销降两个数量级。
- 支持无训练在线防护,适合实际部署中对安全与效率要求高的场景。
现代开放世界智能体如OpenClaw具备跨环境执行能力,但引入广泛的新安全风险。同时,前沿大模型显著降低攻击门槛,现有对齐框架难以满足真实部署需求。为此,我们提出轻量级可扩展的智能体安全对齐框架AgentDoG 1.5。具体而言,更新智能体安全分类体系以涵盖Codex和OpenClaw执行场景中的新兴风险;构建基于分类体系的数据引擎,并结合影响函数净化技术,仅用约1000个样本即训练出0.8B、2B、4B和8B参数的AgentDoG 1.5变体,其性能可比肩领先闭源模型(如GPT-5.4)。基于该框架,我们搭建高效智能体安全SFT与RL训练环境,使Docker级部署开销降低两个数量级。最终,将AgentDoG 1.5作为无需训练的在线防护屏障,实现实时安全管控。大量实验表明,该框架在多样复杂交互场景中达到顶尖表现。所有模型与数据集均开源。
原文摘要 · Abstract (English)
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI models drastically lower attack barriers, rendering current agent alignment frameworks inadequate for real-world deployment. To tackle these emerging threats, we propose a lightweight and scalable agent safety alignment framework. Specifically, we update the agent safety taxonomy to accommodate emergent risks from Codex and OpenClaw execution scenarios. We further build a taxonomy-guided data engine with influence-function purification to train lightweight AgentDoG 1.5 variants (0.8B, 2B, 4B, and 8B parameters) using only around 1k samples, achieving comparable performance with leading closed-source models (e.g., GPT-5.4). Based on AgentDoG 1.5, we construct a highly efficient agentic safety SFT and RL training environment, which reduces deployment overhead in Docker-level environments by two orders of magnitude. Finally, we deploy AgentDoG 1.5 as a training-free online guardrail for real-time safety moderation. Extensive experimental results indicate that AgentDoG 1.5 achieves state-of-the-art performance in diverse and complex interactive agentic scenarios. All models and datasets are openly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。