通过多轮对抗训练提升大模型对话安全性。
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
- 用多轮红队攻击学习恶意对话模式。
- 目标模型在安全基准上性能显著提升。
- 适合关注大模型安全对齐的研究者。
大语言模型(LLMs)面临的越狱攻击日益增多,尤其在多轮对话中,恶意意图可能隐匿于交互过程,导致模型生成有害内容。本文提出多轮安全对齐框架(MTSA),包含两个阶段:在思维引导的攻击学习阶段,红队模型学习生成基于思维链的多轮越狱提示;在对抗迭代优化阶段,红队模型与目标模型持续相互提升交互能力。此外,引入基于未来奖励的多轮强化学习算法,增强安全对齐的鲁棒性。实验表明,红队模型具备当前最优的攻击能力,目标模型在多个安全基准上显著提升表现。
原文摘要 · Abstract (English)
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the \textbf{M}ulti-\textbf{T}urn \textbf{S}afety \textbf{A}lignment (\ourapproach) framework, to address the challenge of securing LLMs in multi-round interactions. It consists of two stages: In the thought-guided attack learning stage, the red-team model learns about thought-guided multi-round jailbreak attacks to generate adversarial prompts. In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction. Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment. Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。