arXiv:2604.20704cs.CRcs.LG2026-04

构建首个结构化文献分析与自动化对抗鲁棒性测试框架,解决评估碎片化问题。

Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing

  • 整合9篇文献与7种协议,系统梳理领域共识与未解难题
  • 实测显示92%梯度掩蔽案例被预筛识别,RDI评分与完整攻击高度相关
  • 支持多范数评估,适配标准合规要求,助力可信AI部署

对抗鲁棒性评估是可信机器学习部署的基础,但当前领域存在协议分散和梯度掩蔽未被发现的问题。本文贡献两点:(1) 结构化综述。通过七种互补协议分析2020—2026年九篇同行评审文献,首次完成该领域共识与未解决问题的端到端结构化分析。(2) Auto-ART框架。提出开源框架,集成50+攻击、28个防御模块、鲁棒性诊断指数(RDI)及梯度掩蔽检测,支持l1/l2/linf/语义/空间多范数评估,并可映射至NIST AI RMF、OWASP LLM Top 10与欧盟人工智能法案。在RobustBench上的实证表明,Auto-ART预筛可识别92%的梯度掩蔽案例,且RDI排名与全量AutoAttack高度一致;多范数评估揭示顶尖模型平均与最差情况间存在23.5个百分点的鲁棒性差距。无先前工作能将结构化元科学研究与可执行评估框架结合,弥合文献空白与工程落地之间的鸿沟。

原文摘要 · Abstract (English)

Adversarial robustness evaluation underpins every claim of trustworthy ML deployment, yet the field suffers from fragmented protocols and undetected gradient masking. We make two contributions. (1) Structured synthesis. We analyze nine peer-reviewed corpus sources (2020--2026) through seven complementary protocols, producing the first end-to-end structured analysis of the field's consensus and unresolved challenges. (2) Auto-ART framework. We introduce Auto-ART, an open-source framework that operationalizes identified gaps: 50+ attacks, 28 defense modules, the Robustness Diagnostic Index (RDI), and gradient-masking detection. It supports multi-norm evaluation (l1/l2/linf/semantic/spatial) and compliance mapping to NIST AI RMF, OWASP LLM Top 10, and the EU AI Act. Empirical validation on RobustBench demonstrates that Auto-ART's pre-screening identifies gradient masking in 92% of flagged cases, and RDI rankings correlate highly with full AutoAttack. Multi-norm evaluation exposes a 23.5 pp gap between average and worst-case robustness on state-of-the-art models. No prior work combines such structured meta-scientific analysis with an executable evaluation framework bridging literature gaps into engineering.

对抗鲁棒性文献综述自动测试可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。