arXiv:2510.26981cs.LGcs.AI2025-10

在计算资源有限时,通过精细调控激活重算提升对抗攻击强度。

Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget

  • 按迭代与层级精细控制激活重算,优化计算分配。
  • 同等预算下攻击效果优于现有方法,攻击成功率显著提升。
  • 适合对抗训练场景,可节省70%计算成本仍保持性能。

本文解决在计算资源受限条件下如何最大化迭代式对抗攻击强度的关键挑战。粗略减少攻击迭代次数虽降低开销,但严重削弱攻击效果。为在预算约束内实现最佳攻击效能,我们提出一种细粒度控制机制,可在迭代层面和层间层面选择性地重新计算激活值。大量实验表明,该方法在相同计算成本下始终优于现有基线。此外,将其集成到对抗训练中,仅需原预算的30%即可达到相近性能。

原文摘要 · Abstract (English)

This work tackles a critical challenge in AI safety research under limited compute: given a fixed computation budget, how can one maximize the strength of iterative adversarial attacks? Coarsely reducing the number of attack iterations lowers cost but substantially weakens effectiveness. To fulfill the attainable attack efficacy within a constrained budget, we propose a fine-grained control mechanism that selectively recomputes layer activations across both iteration-wise and layer-wise levels. Extensive experiments show that our method consistently outperforms existing baselines at equal cost. Moreover, when integrated into adversarial training, it attains comparable performance with only 30% of the original budget.

对抗攻击计算效率模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。