arXiv:2602.07470cs.AI2026-02被引 7

测试大模型推理链在干扰下的鲁棒性,发现模型能恢复但代价是推理变长或变短。

Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?

  • 通过七种干预手段,在固定步骤扰动推理链
  • 模型对多数干扰有鲁棒性,大模型更强,早期干扰更致命
  • 怀疑表达是恢复关键,但恢复会增加长度或降低准确率

推理大模型(RLLMs)在回答复杂任务前生成逐步推理链(CoT),提升性能并增强可解释性。但这些推理链对内部干扰有多强的鲁棒性?我们设计了一个受控评估框架,对多个开源权重的RLLMs在数学、科学和逻辑任务中,于特定时间步施加七种干预(良性、中性、对抗性)。结果表明,RLLMs总体具备鲁棒性,能有效恢复多种扰动,且鲁棒性随模型规模提升,早期干预导致性能下降。然而,鲁棒性不具风格不变性:改写会抑制怀疑表达并降低性能,其他干预则引发怀疑并促进恢复。恢复也有代价:中性和对抗性噪声使推理链长度增加超过200%,而改写虽缩短链条却损害准确性。研究揭示了推理完整性维持机制,确认怀疑是核心恢复信号,并指出未来训练需权衡鲁棒性与效率。

原文摘要 · Abstract (English)

Reasoning LLMs (RLLMs) generate step-by-step chains of thought (CoTs) before giving an answer, which improves performance on complex tasks and makes reasoning more transparent. But how robust are these reasoning traces to disruptions that occur within them? To address this question, we introduce a controlled evaluation framework that perturbs a model's own CoT at fixed timesteps. We design seven interventions (benign, neutral, and adversarial) and apply them to multiple open-weight RLLMs across Math, Science, and Logic tasks. Our results show that RLLMs are generally robust, reliably recovering from diverse perturbations, with robustness improving with model size and degrading when interventions occur early. However, robustness is not style-invariant: paraphrasing suppresses doubt-like expressions and reduces performance, while other interventions trigger doubt and support recovery. Recovery also carries a cost: neutral and adversarial noise can inflate CoT length by more than 200%, whereas paraphrasing shortens traces but harms accuracy. These findings provide new evidence on how RLLMs maintain reasoning integrity, identify doubt as a central recovery mechanism, and highlight trade-offs between robustness and efficiency that future training methods should address.

推理链鲁棒性大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。