arXiv:2502.02260cs.LGcs.CR2025-02中稿 · ICML被引 24

LLM对抗攻击研究进展缓慢,问题定义不清、求解困难且评估不严谨。

Position: Adversarial ML for LLMs Is Not Making Any Progress

  • 指出当前对抗机器学习在大模型时代面临问题定义模糊、求解难度高
  • 强调评估方法缺乏严谨性,阻碍了真正进展
  • 提醒学界警惕重复投入却无实质突破的风险

过去十年,大量研究致力于保护在对抗环境下的机器学习模型,但即便对于简单的'玩具'问题(如对微小对抗扰动的鲁棒性),进展也十分缓慢,且常因评估不严谨而受阻。如今,对抗机器学习的研究重心已转向更大的通用语言模型。本文认为,当前情况更糟:在大模型时代,该领域研究的问题(1)定义更模糊,(2)更难求解,(3)评估也更为困难。因此我们警告,未来另一个十年的研究可能仍无法带来实质性进步。

原文摘要 · Abstract (English)

In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even for simple "toy" problems (e.g., robustness to small adversarial perturbations) and is often hindered by non-rigorous evaluations. Today, adversarial ML research has shifted towards studying larger, general-purpose language models. In this position paper, we argue that the situation is now even worse: in the era of LLMs, the field of adversarial ML studies problems that are (1) less clearly defined, (2) harder to solve, and (3) even more challenging to evaluate. As a result, we caution that yet another decade of work on adversarial ML may be failing to produce meaningful progress.

对抗机器学习大模型安全评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。