arXiv:2506.13726cs.AIcs.CR2025-06中稿 · LLMSEC 2025被引 6

对比推理模型与普通模型的抗攻击能力,发现其安全表现因攻击类型而异。

Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

  • 系统测试多种提示攻击,比较推理模型与非推理模型的安全性差异。
  • 平均攻击成功率:推理模型42.51%,非推理模型45.53%,略更稳健。
  • 部分攻击下推理模型反而更脆弱,如树状攻击中差达32个百分点。

先进推理能力虽提升了大语言模型在数学与编程基准上的表现,但其对对抗性提示攻击的脆弱性尚不明确。本文系统评估了推理增强模型与同类非推理模型在多种提示攻击类别下的安全性。实验结果显示,总体上推理模型攻击成功率(42.51%)低于非推理模型(45.53%),略有优势。然而,这一整体趋势掩盖了显著的类别差异:某些攻击类型下,推理模型明显更脆弱(如树状攻击中高达32个百分点恶化),而在其他类型中则显著更鲁棒(如跨站脚本注入攻击中提升29.8个百分点)。研究揭示了推理能力对模型安全性的复杂影响,强调需针对多样化对抗技术开展压力测试。

原文摘要 · Abstract (English)

The introduction of advanced reasoning capabilities have improved the problem-solving performance of large language models, particularly on math and coding benchmarks. However, it remains unclear whether these reasoning models are more or less vulnerable to adversarial prompt attacks than their non-reasoning counterparts. In this work, we present a systematic evaluation of weaknesses in advanced reasoning models compared to similar non-reasoning models across a diverse set of prompt-based attack categories. Using experimental data, we find that on average the reasoning-augmented models are \emph{slightly more robust} than non-reasoning models (42.51\% vs 45.53\% attack success rate, lower is better). However, this overall trend masks significant category-specific differences: for certain attack types the reasoning models are substantially \emph{more vulnerable} (e.g., up to 32 percentage points worse on a tree-of-attacks prompt), while for others they are markedly \emph{more robust} (e.g., 29.8 points better on cross-site scripting injection). Our findings highlight the nuanced security implications of advanced reasoning in language models and emphasize the importance of stress-testing safety across diverse adversarial techniques.

模型安全对抗攻击推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。