揭示对话中模型安全失效的动态演化机制,发现看似安全的模型在连贯对话中会突然失守。
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- 将对话历史视为状态转移,用可控方式分析模型安全边界跨越过程。
- 多模型测试显示,静态评估下安全的模型在连续对话中会快速崩溃。
- 适合关注模型安全性、对抗攻击与对话系统设计的研究者。
大型语言模型的安全对齐通常在孤立查询下评估,但真实使用场景具有多轮对话特性。尽管多轮越狱攻击已证明有效,其背后的对话安全失效结构仍不明确。本文从状态空间视角研究安全失败,发现当前安全对齐模型中的许多多轮安全问题源于上下文状态的演化,这是孤立提示分析无法完全捕捉的。我们提出STAR框架,将对话历史视为状态转移算子,实现对交互轨迹上安全行为的受控分析。不同于优化攻击强度,STAR提供一种原则性探针,考察对齐模型在自回归条件下的安全边界穿越过程。在多个前沿语言模型中,我们发现即使在静态评估下表现稳健的系统,在有结构的多轮交互中也会出现快速且可复现的安全崩溃。机制分析揭示了拒绝相关表征的单调漂移以及由角色条件上下文引发的突变相变。这些发现表明,语言模型安全应被视为定义于对话轨迹上的动态、状态依赖过程。
原文摘要 · Abstract (English)
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure remains insufficiently understood. In this work, we study safety failures from a state-space perspective and show that many multi-turn safety failures in current safety-aligned language models arise from contextual state evolution, a regime that is not fully captured by isolated prompt-level analyses alone. We introduce STAR, a state-oriented diagnostic framework that treats dialogue history as a state transition operator and enables controlled analysis of safety behavior along interaction trajectories. Rather than optimizing attack strength, STAR provides a principled probe of how aligned models traverse the safety boundary under autoregressive conditioning. Across multiple frontier language models, we find that systems that appear robust under static evaluation can undergo rapid and reproducible safety collapse under structured multi-turn interaction. Mechanistic analysis reveals monotonic drift away from refusal-related representations and abrupt phase transitions induced by role-conditioned context. Together, these findings motivate viewing language model safety as a dynamic, state-dependent process defined over conversational trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。