揭示角色扮演攻击如何绕过模型安全机制,定位关键漏洞。
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

- 通过对比实验分析角色扮演包装对拒绝行为的影响。
- 发现有害请求的识别信号在回答阶段减弱,形成安全传递衰减。
- 提出修复方向:强化危害识别到拒绝之间的内在连接。
大型语言模型在遵循指令的同时会拒绝有害请求。越狱攻击利用这一平衡诱导模型生成本应拒绝的内容。角色扮演越狱尤为严重:有害请求可隐藏于人物、场景和任务构成的角色扮演包装中,但模型仍可能响应。我们采用机制可解释性方法,探究上下文如何逆转拒绝行为及其贡献组件。在两个基准测试、三个模型家族和四个自定义包装下,比较带有与不带包装的有害与良性请求。追踪从请求到最终提示状态的隐状态差异,通过受控反事实隔离包装操作,干预其激活方向,并几何分解有效方向。分析得出三项发现:(1) 成功攻击在请求阶段仍保留有害与良性区分,但在回答起始阶段拒绝相关表达减弱,此现象称为安全传递衰减;(2) 在请求周围构建完整角色扮演并置于场景框架内具有因果作用:移除相关激活变化后恢复拒绝能力;(3) 这些效应主要共享内部结构,多数修复可通过与模型常规拒绝机制对齐的组件实现,而场景框架仅保留较小、依赖模型的成分。研究解释了角色扮演为何能导致合规,同时指明未来防护的明确目标:保持从危害识别到拒绝的连通性。
原文摘要 · Abstract (English)
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmarks, three model families, and four authored wrappers, we compare matched harmful and benign requests with and without this wrapper. We trace hidden-state contrasts from the request to the final prompt state, isolate wrapper operations through controlled counterfactuals, intervene on their activation directions in held-out evaluation requests, and decompose effective directions geometrically. Our analysis yields three findings. (1) Successful attacks retain the measured harmful-versus-benign distinction at the request, while its refusal-associated expression weakens where the answer begins, a pattern we call safety-relay attenuation. (2) Constructing the complete roleplay around the request and framing it within the scenario contribute causally: removing the associated activation changes restores refusal. (3) These effects largely share internal structure, and most repair is reproduced by components aligned with the model's ordinary refusal of harmful requests without roleplay; scenario framing retains a smaller, model-dependent component. Together, these findings explain how roleplay can produce compliance despite retained evidence of harm and identify a concrete target for future safeguards: maintaining the connection from harm recognition to refusal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。