解决角色扮演对话中安全与实用性难以兼顾的难题
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents
- 根据风险耦合程度动态调整安全与实用偏好
- 在保持角色表现力的同时,显著提升内容安全性
- 适合研究对话系统安全机制或角色扮演应用的开发者
大型语言模型在角色扮演对话代理方面取得显著进展,展现出角色模拟的实用价值。然而,如何在角色表现力与内容安全之间取得平衡仍具挑战性,因为角色模拟常伴随生成不安全内容的风险。本文首次系统分析了多款LLM中的安全-实用性权衡问题,发现反派角色与用户提问之间的风险耦合是导致该权衡的关键因素。基于此,提出自适应动态多偏好(ADMP)方法,根据风险耦合程度动态调节模型输出的安全与实用性倾向,并引入耦合边缘采样(CMS)以增强高风险场景下的检测能力。实验表明,该方法在维持角色表现力的同时,有效提升了安全性指标。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made remarkable advances in role-playing dialogue agents, demonstrating their utility in character simulations. However, it remains challenging for these agents to balance character portrayal utility with content safety because this essential character simulation often comes with the risk of generating unsafe content. To address this issue, we first conduct a systematic exploration of the safety-utility trade-off across multiple LLMs. Our analysis reveals that risk scenarios created by villain characters and user queries (referred to as risk coupling) contribute to this trade-off. Building on this, we propose a novel Adaptive Dynamic Multi-Preference (ADMP) method, which dynamically adjusts safety-utility preferences based on the degree of risk coupling and guides the model to generate responses biased toward utility or safety. We further introduce Coupling Margin Sampling (CMS) into coupling detection to enhance the model's ability to handle high-risk scenarios. Experimental results demonstrate that our approach improves safety metrics while maintaining utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。