arXiv:2603.04913cs.ROcs.CV2026-03中稿 · ICRA

用3D对抗纹理让机器人误操作,且在不同视角下仍有效。

Beyond the Patch: Exploring Vulnerabilities of Visuomotor Policies via Viewpoint-Consistent 3D Adversarial Object

  • 通过可微渲染优化3D物体的对抗纹理,适应多视角变化。
  • 在多种环境和距离下使机器人误触目标,黑盒迁移成功率高。
  • 适合研究机器人安全、视觉攻击与防御的学者参考。

基于神经网络的视觉-运动策略虽能完成操作任务,但易受感知攻击影响。传统2D对抗补丁在固定摄像头下有效,但在动态视角(如腕戴相机)下因透视畸变而失效。为此,本文提出一种基于可微渲染的3D物体视角一致对抗纹理优化方法。采用期望变换(EOT)与粗到精(C2F)训练策略,利用距离相关的频率特性,生成在不同相机-物体距离下均有效的纹理。进一步引入显著性引导扰动以转移策略注意力,并设计目标损失持续引导机器人靠近对抗物体。大量实验表明,该方法在多种环境条件下均有效,且具备黑盒迁移能力与真实场景适用性。

原文摘要 · Abstract (English)

Neural network-based visuomotor policies enable robots to perform manipulation tasks but remain susceptible to perceptual attacks. For example, conventional 2D adversarial patches are effective under fixed-camera setups, where appearance is relatively consistent; however, their efficacy often diminishes under dynamic viewpoints from moving cameras, such as wrist-mounted setups, due to perspective distortions. To proactively investigate potential vulnerabilities beyond 2D patches, this work proposes a viewpoint-consistent adversarial texture optimization method for 3D objects through differentiable rendering. As optimization strategies, we employ Expectation over Transformation (EOT) with a Coarse-to-Fine (C2F) curriculum, exploiting distance-dependent frequency characteristics to induce textures effective across varying camera-object distances. We further integrate saliency-guided perturbations to redirect policy attention and design a targeted loss that persistently drives robots toward adversarial objects. Our comprehensive experiments show that the proposed method is effective under various environmental conditions, while confirming its black-box transferability and real-world applicability.

机器人安全对抗攻击3D生成视觉-运动策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。