提出自动化攻击搜索框架,提升世界模型代理的对抗鲁棒性评估效率与准确性。
WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents

- 将对抗攻击评估转化为有限预算下的配置搜索,融合攻击类型与优化参数。
- 在Atari和DMC任务中使奖励下降率提升至1.034和0.682,显著超越基线。
- 适用于需高效评估世界模型鲁棒性的研究者或安全验证场景。
尽管世界模型作为决策代理日益普及,其对抗鲁棒性仍因缺乏专用自动化评估方法而未被充分探索。主要挑战在于攻击评估需兼顾准确与高效:手动调参攻击会高估鲁棒性,而穷尽超参数搜索因每次候选需通过学习的隐空间动态进行闭环推演而代价过高。本文提出WMAttack,一种面向世界模型代理的自动化攻击搜索框架。该框架将鲁棒性评估建模为对攻击配置(包括攻击类型、扰动预算、优化步数、重启次数与分配规则)的有限预算搜索。为提升搜索精度,引入自校正攻击搜索(SCAS),利用奖励衰减、动作不稳定性、运行成本与推演变异性的反馈来优化攻击提议分布。为提升效率,提出表征引导攻击检索(RGAR),从表征相似任务的历史有效配置中检索,为新环境提供热启动。理论分析表明,当提议分布向高价值攻击集中时,有限预算搜索性能得以提升。在Atari和DeepMind Control任务上,WMAttack始终发现比基线更强的攻击,使DreamerV3 Atari的归一化奖励下降率从0.497提升至1.034,DMC任务从0.319提升至0.682。消融实验进一步表明,RGAR提升初始候选质量,而SCAS在固定预算下提升最终攻击效用。
原文摘要 · Abstract (English)
Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated automated evaluation methods. A key obstacle is that attack evaluation must be both accurate and efficient: weak manually tuned attacks can overestimate robustness, while exhaustive hyperparameter search is prohibitively expensive because each candidate requires closed-loop rollouts through learned latent dynamics. We introduce WMAttack, an automated attack-search framework for adversarial evaluation of world-model agents. WMAttack formulates robustness evaluation as a finite-budget search over attack configurations, including attack families, perturbation budgets, optimization steps, restarts, and allocation rules. To improve search accuracy, Self-Correcting Attack Search (SCAS) refines the attack proposal distribution using feedback from reward degradation, action instability, runtime cost, and rollout variability. To improve search efficiency, Representation-Guided Attack Retrieval (RGAR) retrieves effective historical configurations from representation-similar tasks, providing a warm start for unseen environments. We provide a theoretical explanation showing that proposal refinement improves finite-budget search when it shifts probability mass toward high-utility attacks. Across Atari and DeepMind Control tasks, WMAttack consistently discovers stronger attacks than the evaluated baselines, improving normalized reward drop from 0.497 to 1.034 on DreamerV3 Atari and from 0.319 to 0.682 on DMC. Ablations further show that RGAR improves initial candidate quality and SCAS improves final attack utility under fixed evaluation budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。