arXiv:2507.12314cs.LGcs.AI2025-07被引 3

提出防御思维链攻击的新框架,让大模型自动识别并修复被篡改的推理过程。

Thought Purity: A Defense Framework For Chain-of-Thought Attack

  • 用强化学习训练模型动态识别恶意推理逻辑
  • 在多个模型上将攻击成功率显著降低,且正常任务性能不降
  • 适合关注大模型安全性的研究人员和应用开发者

大型推理模型(LRMs)通过思维链(CoT)推理解决复杂任务,但这一显式推理过程存在关键漏洞:攻击者可操控思维链,引发思维链攻击(CoTA)。此类攻击会微妙地扭曲推理路径,导致错误输出,传统防御方法常以牺牲模型实用性为代价。为此,我们提出思想纯净性(Thought Purity, TP)防御框架,从被动拒绝对抗转向主动推理恢复。TP结合安全感知数据流水线与强化学习,采用双重奖励机制,教会模型动态识别并隔离恶意逻辑,同时保留正确推理能力。多模型家族实验表明,TP显著降低CoTA攻击成功率,且在良性任务上的表现维持或提升。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) leverage Chain-of-Thought (CoT) reasoning to solve complex tasks, but this explicit reasoning process introduces a critical vulnerability: adversarial manipulation of the thought chain itself, known as Chain-of-Thought Attacks (CoTA). Such attacks subtly corrupt the reasoning path to produce erroneous outputs, challenging conventional defenses that often sacrifice model utility for safety. To address this, we propose Thought Purity(TP), a defense framework that shifts from passive refusal to active reasoning recovery. TP integrates a safety-aware data pipeline with reinforcement learning, employing a dual-reward mechanism to teach models to dynamically identify and isolate malicious logic while preserving correct reasoning. Experiments on multiple model families demonstrate that TP significantly reduces the attack success rate of CoTA while maintaining or enhancing the model's performance on benign tasks.

大模型安全思维链对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。