发现黑盒攻击成功率随集成模型数对数线性增长,揭示了攻击的可扩展本质。
Scaling Laws for Black box Adversarial Attacks
- 通过先进优化器缓解梯度冲突,实现大规模模型集成攻击。
- 攻击成功率随集成规模对数线性提升,最大达80%以上。
- 适用于评测大模型鲁棒性,尤其适合评估GPT-4o、Claude-3.5-Sonnet等模型。
对抗样本具有跨模型迁移性,可对商用模型实施威胁性黑盒攻击。模型集成(攻击多个代理模型)是提升迁移性的已知策略,但以往研究多采用小规模固定集成,未探究扩大集成规模的效果。本文首次开展大规模实证研究,通过解决梯度冲突并结合理论分析与实验验证,发现攻击成功率(ASR)与集成规模 $T$ 的对数呈稳健且普遍的线性关系。该规律在标准分类器、前沿防御机制及大语言模型(MLLMs)中均被严格验证,并表明规模化攻击能提炼目标类别的鲁棒语义特征。据此我们对SOTA MLLMs进行基准测试,结果显示对GPT-4o等专有模型攻击成功率超80%,同时凸显Claude-3.5-Sonnet的显著抗性。研究呼吁将鲁棒性评估重点从复杂算法转向对规模化攻击的系统性理解。
原文摘要 · Abstract (English)
Adversarial examples exhibit cross-model transferability, enabling threatening black-box attacks on commercial models. Model ensembling, which attacks multiple surrogate models, is a known strategy to improve this transferability. However, prior studies typically use small, fixed ensembles, which leaves open an intriguing question of whether scaling the number of surrogate models can further improve black-box attacks. In this work, we conduct the first large-scale empirical study of this question. We show that by resolving gradient conflict with advanced optimizers, we discover a robust and universal log-linear scaling law through both theoretical analysis and empirical evaluations: the Attack Success Rate (ASR) scales linearly with the logarithm of the ensemble size $T$. We rigorously verify this law across standard classifiers, SOTA defenses, and MLLMs, and find that scaling distills robust, semantic features of the target class. Consequently, we apply this fundamental insight to benchmark SOTA MLLMs. This reveals both the attack's devastating power and a clear robustness hierarchy: we achieve 80\%+ transfer attack success rate on proprietary models like GPT-4o, while also highlighting the exceptional resilience of Claude-3.5-Sonnet. Our findings urge a shift in focus for robustness evaluation: from designing intricate algorithms on small ensembles to understanding the principled and powerful threat of scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。