发现对话大模型会无休止延长对话,因内在机制诱使模型反复追问。
Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs
- 找到一种通用激活子空间,让模型在任何任务中都倾向追问澄清。
- 攻击可跨提示和任务持续生效,使对话轮次增加3倍以上且仍合规。
- 现有防御措施无效,适合研究模型安全与对抗攻击的读者关注。
多轮对话长度是对话类大模型运行成本的主要因素。本文揭示了一种新型失效模式:轮次放大,即模型在未完成任务时持续延长对话。我们发现,攻击者可系统性利用常见于多轮对话中的澄清请求行为,实现可扩展的对话延长。超越提示层面的行为,我们从机制角度识别出一个与查询无关、普遍存在的激活子空间,对应澄清型回应。不同于依赖逐轮提示优化的以往攻击,该攻击源于对话动态,可在不同提示和任务间持续存在。我们证明,这一机制提供了一条可扩展的轮次放大路径:通过微调供应链攻击或低级参数篡改,在运行时均能一致地使模型转向抽象化、追问式行为。在多个指令微调的大模型及基准测试中,该攻击显著提升对话轮次,同时保持合规性。此外,现有防御手段对此类新缺陷保护有限。
原文摘要 · Abstract (English)
Multi-turn interaction length is a dominant factor in the operational costs of conversational LLMs. In this work, we present a new failure mode in conversational LLMs: turn amplification, in which a model consistently prolongs multi-turn interactions without completing the underlying task. We show that an adversary can systematically exploit clarification-seeking behavior$-$commonly encouraged in multi-turn conversation settings$-$to scalably prolong interactions. Moving beyond prompt-level behaviors, we take a mechanistic perspective and identify a query-independent, universal activation subspace associated with clarification-seeking responses. Unlike prior cost-amplification attacks that rely on per-turn prompt optimization, our attack arises from conversational dynamics and persists across prompts and tasks. We show that this mechanism provides a scalable pathway to induce turn amplification: both supply-chain attacks via fine-tuning and runtime attacks through low-level parameter corruptions consistently shift models toward abstract, clarification-seeking behavior across prompts. Across multiple instruction-tuned LLMs and benchmarks, our attack substantially increases turn count while remaining compliant. We also show that existing defenses offer limited protection against this emerging class of failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。