提出角色幻觉攻击框架,揭示大模型角色偏离根源并设计新防御策略。
RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
- 从攻击角度分析角色幻觉,发现查询稀疏与角色冲突是主因。
- 构建新数据集验证,发现强化模型仍易被攻破。
- 提出叙述模式防御,通过补充叙事提升角色一致性。
由大语言模型驱动的角色扮演系统在情感交互应用中日益重要,但存在角色幻觉问题——模型偏离预设角色生成不一致回应。本文首次从攻击视角系统分析该问题,提出RoleBreak框架,识别出查询稀疏和角色-查询冲突为关键驱动机制。基于此,我们构建了新评估数据集RoleBreakEval,用于测试现有缓解技术。实验表明,即使经过优化的模型仍易受攻击。为此,我们提出新型防御策略Narrator Mode,通过生成叙述性补充上下文来缓解角色-查询冲突,提升查询泛化能力。结果表明,该方法显著优于传统拒绝策略,在降低幻觉、增强角色忠实度与叙事连贯性方面表现更优。
原文摘要 · Abstract (English)
Role-playing systems powered by large language models (LLMs) have become increasingly influential in emotional communication applications. However, these systems are susceptible to character hallucinations, where the model deviates from predefined character roles and generates responses that are inconsistent with the intended persona. This paper presents the first systematic analysis of character hallucination from an attack perspective, introducing the RoleBreak framework. Our framework identifies two core mechanisms-query sparsity and role-query conflict-as key factors driving character hallucination. Leveraging these insights, we construct a novel dataset, RoleBreakEval, to evaluate existing hallucination mitigation techniques. Our experiments reveal that even enhanced models trained to minimize hallucination remain vulnerable to attacks. To address these vulnerabilities, we propose a novel defence strategy, the Narrator Mode, which generates supplemental context through narration to mitigate role-query conflicts and improve query generalization. Experimental results demonstrate that Narrator Mode significantly outperforms traditional refusal-based strategies by reducing hallucinations, enhancing fidelity to character roles and queries, and improving overall narrative coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。