用多个智能体协作防御视觉语言模型的越狱攻击。
Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- 设计四类智能体协同检测与拦截越狱攻击。
- 在5个数据集上降低7-19%攻击成功率,拒绝率提升4-20%。
- 适合关注多模态安全、可解释性防御的研究者。
智能体方法作为一种自主推理与协作的新范式,被拓展至安全对齐领域,提出Agentic Moderation框架。该框架不依赖特定模型,通过盾牌(Shield)、响应(Responder)、评估(Evaluator)和反思(Reflector)四类专用智能体,实现动态、协作式的上下文感知安全防护。相比传统静态分类方法,本方法在五个数据集和四种主流大型视觉语言模型(LVLMs)上验证,使攻击成功率(ASR)降低7-19%,非遵循率(NF)稳定,拒绝率(RR)提升4-20%,实现更鲁棒、可解释且平衡的安全表现。该架构具备模块化、可扩展、细粒度等优势,凸显智能体系统在自动化安全治理中的潜力。
原文摘要 · Abstract (English)
Agentic methods have emerged as a powerful and autonomous paradigm that enhances reasoning, collaboration, and adaptive control, enabling systems to coordinate and independently solve complex tasks. We extend this paradigm to safety alignment by introducing Agentic Moderation, a model-agnostic framework that leverages specialised agents to defend multimodal systems against jailbreak attacks. Unlike prior approaches that apply as a static layer over inputs or outputs and provide only binary classifications (safe or unsafe), our method integrates dynamic, cooperative agents, including Shield, Responder, Evaluator, and Reflector, to achieve context-aware and interpretable moderation. Extensive experiments across five datasets and four representative Large Vision-Language Models (LVLMs) demonstrate that our approach reduces the Attack Success Rate (ASR) by 7-19%, maintains a stable Non-Following Rate (NF), and improves the Refusal Rate (RR) by 4-20%, achieving robust, interpretable, and well-balanced safety performance. By harnessing the flexibility and reasoning capacity of agentic architectures, Agentic Moderation provides modular, scalable, and fine-grained safety enforcement, highlighting the broader potential of agentic systems as a foundation for automated safety governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。