arXiv:2501.18841cs.LGcs.CR2025-01被引 63

增加推理时计算量可显著提升大模型抗攻击能力

Trading Inference-Time Compute for Adversarial Robustness

  • 通过延长模型推理时间,提升其对抗攻击的鲁棒性
  • 随着计算量增加,攻击成功率趋近于零
  • 无需对抗训练,适合关注模型可靠性的人群

我们在推理模型(特别是 OpenAI o1-preview 和 o1-mini)上研究了增加推理时计算量对其对抗攻击鲁棒性的影响。结果表明,在多种攻击下,推理时计算量越大,模型鲁棒性越强。在多数情况下,随着测试时计算量增加,攻击成功的样本比例趋于零。实验未进行对抗训练,仅通过允许模型在推理时消耗更多计算资源来实现。结果表明,推理时计算量具有提升大语言模型对抗鲁棒性的潜力。我们还探索了针对推理模型的新攻击方式,并分析了计算量无法提升可靠性的场景,推测其原因并提出可能改进方向。

原文摘要 · Abstract (English)

We conduct experiments on the impact of increasing inference-time compute in reasoning models (specifically OpenAI o1-preview and o1-mini) on their robustness to adversarial attacks. We find that across a variety of attacks, increased inference-time compute leads to improved robustness. In many cases (with important exceptions), the fraction of model samples where the attack succeeds tends to zero as the amount of test-time compute grows. We perform no adversarial training for the tasks we study, and we increase inference-time compute by simply allowing the models to spend more compute on reasoning, independently of the form of attack. Our results suggest that inference-time compute has the potential to improve adversarial robustness for Large Language Models. We also explore new attacks directed at reasoning models, as well as settings where inference-time compute does not improve reliability, and speculate on the reasons for these as well as ways to address them.

大模型鲁棒性推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。