arXiv:2601.04736cs.CL2026-01被引 1

提升多轮多模态大模型安全性的高效对齐方法

AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

  • 用多轮对话数据+双目标奖励机制,实现冷启动拒绝与持续对齐
  • 攻击成功率降低10%以上,无害性提升8%,帮助性超13%
  • 适合关注多模态对话安全、需低成本对齐的研究者

多模态大语言模型在交互应用中日益普及,但其在多轮多模态场景下的安全漏洞愈发明显——有害意图可跨轮次逐步构建,安全协议随对话推进逐渐失效。现有基于人类反馈强化学习(RLHF)的对齐方法主要针对单轮视觉问答(VQA)任务,依赖昂贵的人工偏好标注,在对话场景中难以扩展。为此,我们构建了InterSafe-V,一个开源的多模态对话数据集,包含11,270条对话和500个专门设计的拒答VQA样本,通过多模型交互生成,更贴近真实场景,并涵盖特定领域定制的VQA对。基于此数据集,我们提出AM$^3$Safety框架,结合冷启动拒答阶段与基于轮次感知双目标奖励的组相对策略优化(GRPO)微调,全面优化多轮对话中的安全性。在Qwen2.5-VL-7B-Instruct和LLaVA-NeXT-7B上的实验表明,该方法使攻击成功率(ASR)下降超过10%,无害性提升至少8%,帮助性提升超过13%,同时保持模型通用能力。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) are increasingly deployed in interactive applications. However, their safety vulnerabilities become pronounced in multi-turn multi-modal scenarios, where harmful intent can be gradually reconstructed across turns, and security protocols fade into oblivion as the conversation progresses. Existing Reinforcement Learning from Human Feedback (RLHF) alignment methods are largely developed for single-turn visual question-answer (VQA) task and often require costly manual preference annotations, limiting their effectiveness and scalability in dialogues. To address this challenge, we present InterSafe-V, an open-source multi-modal dialogue dataset containing 11,270 dialogues and 500 specially designed refusal VQA samples. This dataset, constructed through interaction between several models, is designed to more accurately reflect real-world scenarios and includes specialized VQA pairs tailored for specific domains. Building on this dataset, we propose AM$^3$Safety, a framework that combines a cold-start refusal phase with Group Relative Policy Optimization (GRPO) fine-tuning using turn-aware dual-objective rewards across entire dialogues. Experiments on Qwen2.5-VL-7B-Instruct and LLaVA-NeXT-7B show more than 10\% decrease in Attack Success Rate (ASR) together with an increment of at least 8\% in harmless dimension and over 13\% in helpful dimension of MLLMs on multi-modal multi-turn safety benchmarks, while preserving their general abilities.

多模态安全对齐方法多轮对话强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。