arXiv:2511.21192cs.CVcs.AI2025-11中稿 · CVPR被引 10

提出可通用迁移的物理攻击贴纸,骗机器人误操作。

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

  • 在共享特征空间中训练单一贴纸,提升跨模型迁移能力。
  • 在多种机器人模型和任务上实现超90%攻击成功率。
  • 适合研究安全漏洞或防御机制的开发者参考。

视觉-语言-动作(VLA)模型易受对抗攻击,但通用且可迁移的攻击仍研究不足,因多数贴纸仅针对单个模型且在黑盒场景下失效。本文系统研究了未知架构、微调版本及仿真到现实迁移下的通用可迁移对抗贴纸攻击。提出UPA-RFAS框架:通过特征空间目标与ℓ₁偏差先验、排斥式InfoNCE损失诱导可迁移表征偏移;采用增强鲁棒性的两阶段极小极大过程,内层学习不可见样本扰动,外层优化通用贴纸以对抗此强化邻域;引入两种VLA特有损失:贴纸注意力主导性以劫持文本→视觉注意力,贴图语义错位性以引发图像-文本不匹配,无需标签。实验涵盖多种VLA模型、操控任务及物理执行,结果表明UPA-RFAS在模型、任务和视角间持续迁移,攻击成功率超90%,揭示了基于贴纸的实际攻击面,并为未来防御建立强基准。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap, we present a systematic study of universal, transferable adversarial patches against VLA-driven robots under unknown architectures, finetuned variants, and sim-to-real shifts. We introduce UPA-RFAS (Universal Patch Attack via Robust Feature, Attention, and Semantics), a unified framework that learns a single physical patch in a shared feature space while promoting cross-model transfer. UPA-RFAS combines (i) a feature-space objective with an $\ell_1$ deviation prior and repulsive InfoNCE loss to induce transferable representation shifts, (ii) a robustness-augmented two-phase min-max procedure where an inner loop learns invisible sample-wise perturbations and an outer loop optimizes the universal patch against this hardened neighborhood, and (iii) two VLA-specific losses: Patch Attention Dominance to hijack text$\to$vision attention and Patch Semantic Misalignment to induce image-text mismatch without labels. Experiments across diverse VLA models, manipulation suites, and physical executions show that UPA-RFAS consistently transfers across models, tasks, and viewpoints, exposing a practical patch-based attack surface and establishing a strong baseline for future defenses.

对抗攻击机器人安全VLA模型物理攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。