用分散对话劫持模型,暴露大模型无状态防护漏洞
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

- 将恶意指令拆分到独立对话中,绕过无状态审核机制
- 测试发现多数主流模型(含OpenAI、Anthropic等)易受攻击
- 适合安全研究人员和模型开发者关注防御新威胁
大型语言模型(LLMs)正被广泛应用于敏感场景,其对抗鲁棒性与安全性愈发关键。本文提出一种名为瞬态轮次注入(Transient Turn Injection, TTI)的新式多轮攻击方法,通过将恶意意图分散至孤立的交互中,系统性地利用无状态内容审核的漏洞。该方法借助由大模型驱动的自动化攻击代理,迭代测试并规避商业及开源模型中的政策约束,突破了传统越狱攻击依赖持续对话上下文的局限。我们在包括OpenAI、Anthropic、Google Gemini、Meta在内的多个前沿模型上开展评估,发现各模型对TTI攻击的抵抗能力差异显著,仅少数架构具备较强内在鲁棒性。我们的自动化黑盒评估框架还揭示了此前未知的模型特异性漏洞与攻击面模式,尤其在医疗等高风险领域表现突出。我们对比了TTI与现有对抗提示方法,并提出会话级上下文聚合与深度对齐等实用缓解策略。研究强调需构建全局、上下文感知的防御体系,并开展持续对抗测试,以应对不断演化的多轮威胁。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This paper introduces Transient Turn Injection(TTI), a new multi-turn attack technique that systematically exploits stateless moderation by distributing adversarial intent across isolated interactions. TTI leverages automated attacker agents powered by large language models to iteratively test and evade policy enforcement in both commercial and open-source LLMs, marking a departure from conventional jailbreak approaches that typically depend on maintaining persistent conversational context. Our extensive evaluation across state-of-the-art models-including those from OpenAI, Anthropic, Google Gemini, Meta, and prominent open-source alternatives-uncovers significant variations in resilience to TTI attacks, with only select architectures exhibiting substantial inherent robustness. Our automated blackbox evaluation framework also uncovers previously unknown model specific vulnerabilities and attack surface patterns, especially within medical and high stakes domains. We further compare TTI against established adversarial prompting methods and detail practical mitigation strategies, such as session level context aggregation and deep alignment approaches. Our study underscores the urgent need for holistic, context aware defenses and continuous adversarial testing to future proof LLM deployments against evolving multi-turn threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。