角色扮演微调会降低大模型安全性能,本文提出新方法平衡角色表现与安全。
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
- 基于角色特性训练95个模型,评估安全风险
- 角色微调后安全性能显著下降,恶角风险更高
- 提出SaRFT方法,兼顾角色扮演与安全性
角色扮演使大语言模型能与用户进行沉浸式个性化互动,但也带来显著安全风险。现有角色扮演微调技术虽提升角色适应性,却可能损害安全性能,尤其对反派角色更为明显。本文首次通过在RoleBench上训练95个角色专属模型,全面评估角色微调的安全风险。实验发现,角色微调导致安全性能明显下降,且风险程度随角色特质变化。为此,我们提出安全感知的角色扮演微调(SaRFT)方法,旨在平衡角色表现与安全性。在LLaMA-3-8B-Instruct、Gemma-2-9B-it和Qwen2.5-7B-Instruct上的大量实验表明,无论采用LoRA还是全参数微调,SaRFT均持续优于当前最优基线。研究强调了角色自适应安全措施的必要性,并为缓解角色特定安全风险提供了实用见解。
原文摘要 · Abstract (English)
Role-playing enables large language models (LLMs) to engage users in immersive and personalized interactions, but it also introduces significant safety risks. Existing role-play fine-tuning techniques improve role adaptability but may degrade safety performance, particularly for villainous characters. In this work, we conduct the first comprehensive assessment of role-play fine-tuning risks by training 95 role-specific LLMs using RoleBench. Our experiments reveal that role-play fine-tuning leads to a noticeable decline in safety performance, with safety risks varying based on character traits. To tackle this challenge, we propose Safety-Aware Role-Play Fine-Tuning (SaRFT), a novel method designed to balance role-playing capabilities and safety. Extensive experiments on LLaMA-3-8B-Instruct, Gemma-2-9B-it, and Qwen2.5-7B-Instruct demonstrate that SaRFT consistently outperforms state-of-the-art baselines under both LoRA and full-parameter fine-tuning settings. Our findings highlight the necessity of role-adaptive safety measures and provide insights into mitigating role-specific safety risks in role-playing LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。