arXiv:2511.09681cs.LGcs.AI2025-11中稿 · CVPR

高效黑盒攻击视觉强化学习,少调环境多骗模型。

SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning

  • 用影子模型+生成对抗网络,模拟攻击效果
  • 在MuJoCo和Atari上减少90%以上环境交互
  • 攻击时图像几乎看不出异常,适合真实场景

视觉强化学习在视觉控制与机器人领域取得显著进展,但其对对抗扰动的脆弱性仍研究不足。现有黑盒攻击多针对向量或离散动作的RL,难以有效作用于图像输入的连续控制任务,主要受限于庞大的动作空间和过多的环境查询。本文提出SEBA,一种面向视觉强化学习代理的样本高效黑盒攻击框架。SEBA集成一个影子Q模型,用于估计对抗条件下的累积奖励;一个生成对抗网络,生成视觉上不可察觉的扰动;以及一个世界模型,模拟环境动态以减少真实环境交互。通过交替训练影子模型与生成器的两阶段迭代过程,SEBA在保持高效的同时实现强攻击性能。在MuJoCo和Atari基准上的实验表明,SEBA显著降低累积奖励,保持视觉保真度,并相比先前黑盒与白盒方法大幅减少环境交互次数。

原文摘要 · Abstract (English)

Visual reinforcement learning has achieved remarkable progress in visual control and robotics, but its vulnerability to adversarial perturbations remains underexplored. Most existing black-box attacks focus on vector-based or discrete-action RL, and their effectiveness on image-based continuous control is limited by the large action space and excessive environment queries. We propose SEBA, a sample-efficient framework for black-box adversarial attacks on visual RL agents. SEBA integrates a shadow Q model that estimates cumulative rewards under adversarial conditions, a generative adversarial network that produces visually imperceptible perturbations, and a world model that simulates environment dynamics to reduce real-world queries. Through a two-stage iterative training procedure that alternates between learning the shadow model and refining the generator, SEBA achieves strong attack performance while maintaining efficiency. Experiments on MuJoCo and Atari benchmarks show that SEBA significantly reduces cumulative rewards, preserves visual fidelity, and greatly decreases environment interactions compared to prior black-box and white-box methods.

对抗攻击强化学习视觉控制高效攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。