无需训练,通过双循环机制让角色扮演模型既忠于人设又抗越狱攻击。
Stay in Character, Stay Safe: Dual-Cycle Adversarial Self-Evolution for Safety Role-Playing Agents
- 用对抗生成与知识提炼双循环,动态提升安全防御能力。
- 在多个闭源大模型上实现角色一致性与越狱抵御力双提升。
- 适合需要高安全性角色扮演的应用,如虚拟客服、教育助手。
基于大语言模型的角色扮演虽在表现力上快速进步,但更强的人设遵循性常导致越狱攻击漏洞加剧,尤其对敏感或负面人设更明显。现有方法多依赖训练阶段的解决方案(如数据筛选或对齐正则化),但维护成本高、易削弱角色一致性,且难以应用于前沿闭源大模型。本文提出一种免训练的双循环对抗自演化框架,包含两个耦合循环:人设定向攻击者循环生成渐进式更强的越狱提示,角色扮演防御者循环将观测到的失败案例提炼为三级知识库——全局安全规则、基于人设的约束条件以及安全合规的角色示例。推理时,防御者从该层级知识库中检索并组合结构化信息以引导生成,使响应在保持目标人设忠实性的同时满足安全要求。大量实验在多个专有大模型上验证了该方法在角色一致性与越狱抵御力上的持续优势,并对未见人设和攻击提示具有强泛化能力。
原文摘要 · Abstract (English)
LLM-based role-playing has rapidly improved in fidelity, yet stronger adherence to persona constraints commonly increases vulnerability to jailbreak attacks, especially for risky or negative personas. Most prior work mitigates this issue with training-time solutions (e.g., data curation or alignment-oriented regularization). However, these approaches are costly to maintain as personas and attack strategies evolve, can degrade in-character behavior, and are typically infeasible for frontier closed-weight LLMs. We propose a training-free Dual-Cycle Adversarial Self-Evolution framework with two coupled cycles. A Persona-Targeted Attacker Cycle synthesizes progressively stronger jailbreak prompts, while a Role-Playing Defender Cycle distills observed failures into a hierarchical knowledge base of (i) global safety rules, (ii) persona-grounded constraints, and (iii) safe in-character exemplars. At inference time, the Defender retrieves and composes structured knowledge from this hierarchy to guide generation, producing responses that remain faithful to the target persona while satisfying safety constraints. Extensive experiments across multiple proprietary LLMs show consistent gains over strong baselines on both role fidelity and jailbreak resistance, and robust generalization to unseen personas and attack prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。