arXiv:2503.19326cs.AI2025-03被引 9

操纵推理链末尾结果可误导大模型忽略正确步骤

Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps

  • 通过篡改推理链末尾计算结果,诱导模型忽视正确过程
  • 多模型测试显示,末尾篡改比结构改动影响更大
  • 发现DeepSeek-R1存在完全中断推理的漏洞,适合安全研究者

近期推理型大语言模型通过长链思维(Chain-of-Thought)显著提升了数学推理能力,其推理标记支持内部自纠正,增强鲁棒性。我们探究此类模型对输入推理链中细微错误的脆弱性,提出‘妥协思维’(CPT)漏洞:当模型面对包含被篡改计算结果的推理标记时,会忽略正确推理步骤而采纳错误结果。通过三种逐步明确的提示方法在多个推理模型上系统评估CPT抗性,发现模型难以识别并纠正此类篡改。值得注意的是,与现有研究认为结构改变比内容修改影响更大的观点相反,我们发现局部末尾标记篡改对推理结果的影响大于结构变化。此外,我们还发现DeepSeek-R1存在漏洞,被篡改的推理标记可导致整个推理过程完全终止。本工作深化了对推理鲁棒性的理解,并揭示了推理密集型应用中的安全风险。

原文摘要 · Abstract (English)

Recent reasoning large language models (LLMs) have demonstrated remarkable improvements in mathematical reasoning capabilities through long Chain-of-Thought. The reasoning tokens of these models enable self-correction within reasoning chains, enhancing robustness. This motivates our exploration: how vulnerable are reasoning LLMs to subtle errors in their input reasoning chains? We introduce "Compromising Thought" (CPT), a vulnerability where models presented with reasoning tokens containing manipulated calculation results tend to ignore correct reasoning steps and adopt incorrect results instead. Through systematic evaluation across multiple reasoning LLMs, we design three increasingly explicit prompting methods to measure CPT resistance, revealing that models struggle significantly to identify and correct these manipulations. Notably, contrary to existing research suggesting structural alterations affect model performance more than content modifications, we find that local ending token manipulations have greater impact on reasoning outcomes than structural changes. Moreover, we discover a security vulnerability in DeepSeek-R1 where tampered reasoning tokens can trigger complete reasoning cessation. Our work enhances understanding of reasoning robustness and highlights security considerations for reasoning-intensive applications.

推理安全大模型漏洞链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。