暗中操控推理过程,实现无需修改输入的隐蔽后门攻击
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
- 在模型内部推理链中植入潜藏触发器,不改用户输入即可发动攻击
- 八组数据集测试中攻击成功率均超90%,跨领域稳定生效
- 适合研究模型安全、对抗攻防的学者与开发者关注
随着个性化AI的快速发展,配备思维链(Chain of Thought, COT)推理能力的定制化大语言模型已广泛应用于数百万AI代理。然而,其复杂的推理过程带来了尚未充分探索的安全隐患。本文提出DarkMind,一种新型的隐式推理层级后门攻击,通过操纵内部思维链步骤,在不改变用户查询的前提下对定制化LLM实施攻击。与以往基于提示的攻击不同,DarkMind利用潜藏触发器在推理链中隐蔽激活,实现恶意行为,且无需修改输入提示或访问模型参数。为确保隐蔽性与可靠性,我们设计了即时与回溯两种触发类型,并整合至统一嵌入模板中以控制触发依赖;采用隐蔽优化算法最小化语义漂移;引入自动化对话起始机制实现跨领域隐蔽激活。在涵盖算术、常识与符号领域的八组推理数据集上,使用五种主流大模型进行的全面实验表明,DarkMind始终维持高攻击成功率。我们进一步探讨防御策略,揭示推理层级后门是一类重大但被忽视的威胁,凸显了构建具备推理感知能力的鲁棒安全机制的迫切需求。
原文摘要 · Abstract (English)
With the rapid rise of personalized AI, customized large language models (LLMs) equipped with Chain of Thought (COT) reasoning now power millions of AI agents. However, their complex reasoning processes introduce new and largely unexplored security vulnerabilities. We present DarkMind, a novel latent reasoning level backdoor attack that targets customized LLMs by manipulating internal COT steps without altering user queries. Unlike prior prompt based attacks, DarkMind activates covertly within the reasoning chain via latent triggers, enabling adversarial behaviors without modifying input prompts or requiring access to model parameters. To achieve stealth and reliability, we propose dual trigger types instant and retrospective and integrate them within a unified embedding template that governs trigger dependent activation, employ a stealth optimization algorithm to minimize semantic drift, and introduce an automated conversation starter for covert activation across domains. Comprehensive experiments on eight reasoning datasets spanning arithmetic, commonsense, and symbolic domains, using five LLMs, demonstrate that DarkMind consistently achieves high attack success rates. We further investigate defense strategies to mitigate these risks and reveal that reasoning level backdoors represent a significant yet underexplored threat, underscoring the need for robust, reasoning aware security mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。