大模型演反派总不靠谱,因安全对齐机制干扰角色真实性
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- 构建四层道德量表评测集,量化评估模型演反派能力
- 越邪恶的角色表现越差,欺骗与操控类特质最难以还原
- 聊天能力强的模型反而演反派更差,适合研究角色扮演安全边界
大型语言模型在创意生成中被越来越多地用于模拟虚构角色,但其在扮演非利他性、对抗性人格方面的能力仍缺乏系统研究。我们假设现代模型的安全对齐机制与真实演绎道德模糊或反派角色存在根本冲突。为此,我们提出了Moral RolePlay基准数据集,包含四层级道德对齐尺度和平衡测试集,用于严格评估。我们让顶尖LLMs扮演从道德楷模到纯粹反派的角色。大规模评估显示,角色道德性越低,扮演真实度越差,尤其在‘欺骗’‘操纵’等与安全原则相悖的特质上表现最弱,常以表面攻击替代深层恶意。此外,通用聊天能力无法预测反派扮演水平,高度安全对齐模型表现尤为不佳。本工作首次系统揭示了该关键局限,凸显安全对齐与创作真实性的张力,为发展更精细、情境感知的对齐方法提供基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of fictional characters. However, their ability to portray non-prosocial, antagonistic personas remains largely unexamined. We hypothesize that the safety alignment of modern LLMs creates a fundamental conflict with the task of authentically role-playing morally ambiguous or villainous characters. To investigate this, we introduce the Moral RolePlay benchmark, a new dataset featuring a four-level moral alignment scale and a balanced test set for rigorous evaluation. We task state-of-the-art LLMs with role-playing characters from moral paragons to pure villains. Our large-scale evaluation reveals a consistent, monotonic decline in role-playing fidelity as character morality decreases. We find that models struggle most with traits directly antithetical to safety principles, such as ``Deceitful'' and ``Manipulative'', often substituting nuanced malevolence with superficial aggression. Furthermore, we demonstrate that general chatbot proficiency is a poor predictor of villain role-playing ability, with highly safety-aligned models performing particularly poorly. Our work provides the first systematic evidence of this critical limitation, highlighting a key tension between model safety and creative fidelity. Our benchmark and findings pave the way for developing more nuanced, context-aware alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。