arXiv:2602.13093cs.AIcs.CL2026-02被引 3

测试九种大模型在多轮攻击下的稳定性,发现推理能力不等于抗攻击能力。

Consistency of Large Reasoning Models Under Multi-Turn Attacks

  • 通过多轮对抗攻击测试,识别出五类失败模式,自怀疑虑和从众心理占一半。
  • 所有模型均易受误导建议影响,社会压力对不同模型效果各异。
  • 传统信心防御失效,随机信心嵌入反而更优,需重新设计防御机制。

具备推理能力的大模型在复杂任务上表现优异,但其在多轮对抗攻击下的鲁棒性仍待探索。我们评估了九种前沿推理模型在对抗攻击下的表现。结果表明,推理虽带来一定但非完整的鲁棒性:多数模型显著优于指令微调基线,但均表现出独特的脆弱性特征——误导建议普遍有效,社会压力则具模型特异性。轨迹分析揭示五类失败模式(自我怀疑、从众倾向、建议劫持、情绪易感、推理疲劳),前两类占失败总数的50%。我们进一步证明,适用于标准大模型的信心感知响应生成(CARG)在推理模型中失效,因长推理链诱发过度自信;反直觉的是,随机信心嵌入表现优于针对性提取。结果表明,推理能力并不自动赋予抗对抗鲁棒性,基于信心的防御需为推理模型重新设计。

原文摘要 · Abstract (English)

Large reasoning models with reasoning capabilities achieve state-of-the-art performance on complex tasks, but their robustness under multi-turn adversarial pressure remains underexplored. We evaluate nine frontier reasoning models under adversarial attacks. Our findings reveal that reasoning confers meaningful but incomplete robustness: most reasoning models studied significantly outperform instruction-tuned baselines, yet all exhibit distinct vulnerability profiles, with misleading suggestions universally effective and social pressure showing model-specific efficacy. Through trajectory analysis, we identify five failure modes (Self-Doubt, Social Conformity, Suggestion Hijacking, Emotional Susceptibility, and Reasoning Fatigue) with the first two accounting for 50% of failures. We further demonstrate that Confidence-Aware Response Generation (CARG), effective for standard LLMs, fails for reasoning models due to overconfidence induced by extended reasoning traces; counterintuitively, random confidence embedding outperforms targeted extraction. Our results highlight that reasoning capabilities do not automatically confer adversarial robustness and that confidence-based defenses require fundamental redesign for reasoning models.

推理模型对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。