arXiv:2512.07059cs.CL2025-12

测试万亿参数大模型在多轮攻击下的安全漏洞,发现模型规模不影响抗攻击能力。

Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models

  • 用多轮对抗攻击框架测评十款前沿模型的安全性。
  • 六款模型攻击成功率96%~100%,四款仅42%~78%。
  • 思考模式推理可将攻击成功率从97%降至42%,适合部署防护。

尽管在安全对齐上投入巨大,大语言模型对复杂多轮对抗攻击的脆弱性仍不明确,模型规模或推理模式是否影响鲁棒性尚不清楚。本研究采用TEMPEST多轮攻击框架,评估了来自八家厂商的十款前沿模型在1000种有害行为上的表现,生成超过97,000次API查询,并通过独立安全分类器自动评估对抗对话效果。结果表明存在显著脆弱性差异:六款模型攻击成功率达96%至100%,四款表现出明显抵抗力,攻击成功率介于42%至78%;相同架构下启用扩展推理模式后,攻击成功率从97%降至42%。研究发现,不同厂商的安全对齐质量差异显著,模型规模无法预测对抗鲁棒性,而思考模式推理提供了一种可部署的安全增强方案。总体而言,当前对齐技术对自适应多轮攻击仍普遍脆弱,但推理时的深思模式是值得探索的防御方向。

原文摘要 · Abstract (English)

Despite substantial investment in safety alignment, the vulnerability of large language models to sophisticated multi-turn adversarial attacks remains poorly characterized, and whether model scale or inference mode affects robustness is unknown. This study employed the TEMPEST multi-turn attack framework to evaluate ten frontier models from eight vendors across 1,000 harmful behaviors, generating over 97,000 API queries across adversarial conversations with automated evaluation by independent safety classifiers. Results demonstrated a spectrum of vulnerability: six models achieved 96% to 100% attack success rate (ASR), while four showed meaningful resistance, with ASR ranging from 42% to 78%; enabling extended reasoning on identical architecture reduced ASR from 97% to 42%. These findings indicate that safety alignment quality varies substantially across vendors, that model scale does not predict adversarial robustness, and that thinking mode provides a deployable safety enhancement. Collectively, this work establishes that current alignment techniques remain fundamentally vulnerable to adaptive multi-turn attacks regardless of model scale, while identifying deliberative inference as a promising defense direction.

对抗攻击大模型安全推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。