通过模拟世界观提升对话安全,揭示多轮越狱攻击的深层机制
Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

- 分离社交影响因子与情境模块,用蒙特卡洛搜索优化多轮策略
- 在六款前沿模型上实现接近极限的越狱成功率,平均仅需2.46次查询
- 发现可执行任务表述能突破拒绝状态,适合安全研究者参考
多轮越狱攻击表明有害意图可分散于对话中,但现有方法未能揭示其驱动机制。本文提出BLUEPRINT框架,将因素化的社交影响策略空间与WORLDVIEWSIM(跨轮情境上下文模块)分离。利用蒙特卡洛树搜索,在四轮对话轨迹中优化18个理论驱动的影响因子组合。在六款前沿模型上,BLUEPRINT实现接近上限的越狱成功率(ASR),平均仅需2.46次查询。结果揭示抵抗型目标存在特定脆弱性:各模型对不同影响因子及策略转换响应各异,但均共享一条共同恢复路径——转向具体、可执行的任务表述能稳定摆脱强硬拒绝状态。消融实验表明:操作性提示影响最大,收益框架异常有效,部分合法性诉求反而适得其反。研究提示,鲁棒多轮安全不仅需监控有害内容,还需关注对话状态如何使不安全请求显得具体且局部可执行。
原文摘要 · Abstract (English)
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。