修复大模型对抗鲁棒性研究的错位目标,推动可测量可复现的进展
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
- 用网络安全分类框架厘清大模型新旧攻击威胁差异
- 指出当前研究因缺乏统一评估导致多数防御无效
- 倡导回归可测量、可复现的学术原则,避免重复历史错误
过去十年中,研究目标错位严重阻碍了对抗鲁棒性研究的进展。例如,过度聚焦优化特定指标而忽视标准化评估,导致研究者采用临时性的启发式防御方法,看似有效实则大多在后续检验中暴露缺陷,对领域进步贡献甚微。本文指出,当前大语言模型(LLMs)的鲁棒性研究正重蹈覆辙,且现实影响可能更严重。为此,我们主张重新校准研究目标以实现真正的对抗对齐。基于已有的网络安全分类体系,我们形式化定义了适用于大语言模型的新型与旧有威胁模型之间的区别。利用该框架,我们强调:要取得进展,必须将对抗对齐拆解为可处理的子问题,并回归可测量性、可复现性和可比性等核心学术原则。尽管挑战巨大,但此次对抗鲁棒性的重新出发,提供了借鉴历史经验、规避旧错误的独特机会。
原文摘要 · Abstract (English)
Misaligned research objectives have considerably hindered progress in adversarial robustness research over the past decade. For instance, an extensive focus on optimizing target metrics, while neglecting rigorous standardized evaluation, has led researchers to pursue ad-hoc heuristic defenses that were seemingly effective. Yet, most of these were exposed as flawed by subsequent evaluations, ultimately contributing little measurable progress to the field. In this position paper, we illustrate that current research on the robustness of large language models (LLMs) risks repeating past patterns with potentially worsened real-world implications. To address this, we argue that realigned objectives are necessary for meaningful progress in adversarial alignment. To this end, we build on established cybersecurity taxonomy to formally define differences between past and emerging threat models that apply to LLMs. Using this framework, we illustrate that progress requires disentangling adversarial alignment into addressable sub-problems and returning to core academic principles, such as measureability, reproducibility, and comparability. Although the field presents significant challenges, the fresh start on adversarial robustness offers the unique opportunity to build on past experience while avoiding previous mistakes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。