arXiv:2604.01618cs.CVcs.AI2026-04被引 2

用可物理部署的3D纹理攻击视觉语言动作模型,让机器人任务失败率达96.7%。

Tex3D: Objects as Attack Surfaces via Adversarial 3D Textures for Vision-Language-Action Models

  • 分离前景背景渲染,实现3D纹理端到端可微优化
  • 在仿真和真实机器人上使任务失败率最高达96.7%
  • 适合研究机器人安全与对抗鲁棒性的学者

视觉语言动作(VLA)模型在机器人操作中表现强劲,但其对真实可实施的对抗攻击的鲁棒性尚未深入探索。现有研究多依赖语言扰动或2D视觉攻击,但这些攻击在现实部署中代表性不足且物理真实性有限。相比之下,对抗性3D纹理更贴近真实场景,因其自然附着于被操作物体,易于在物理环境中部署。然而,将对抗性3D纹理引入VLA系统面临挑战:标准3D模拟器无法从VLA目标函数反向传播至物体外观,难以实现端到端优化。为此,我们提出前景-背景解耦(FBD),通过双渲染器对齐实现可微纹理优化,同时保留原始仿真环境。为进一步确保攻击在长时程及多视角下有效,我们设计轨迹感知对抗优化(TAAO),优先关键帧并采用基于顶点的参数化稳定优化过程。基于上述设计,我们构建了首个在VLA仿真环境中直接进行3D对抗纹理端到端优化的框架Tex3D。仿真与真实机器人实验均表明,Tex3D显著降低多种操作任务的性能,任务失败率最高达96.7%。实证结果揭示了VLA系统对物理可实现的3D对抗攻击存在严重漏洞,凸显了鲁棒性训练的必要性。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have shown strong performance in robotic manipulation, yet their robustness to physically realizable adversarial attacks remains underexplored. Existing studies reveal vulnerabilities through language perturbations and 2D visual attacks, but these attack surfaces are either less representative of real deployment or limited in physical realism. In contrast, adversarial 3D textures pose a more physically plausible and damaging threat, as they are naturally attached to manipulated objects and are easier to deploy in physical environments. Bringing adversarial 3D textures to VLA systems is nevertheless nontrivial. A central obstacle is that standard 3D simulators do not provide a differentiable optimization path from the VLA objective function back to object appearance, making it difficult to optimize through an end-to-end manner. To address this, we introduce Foreground-Background Decoupling (FBD), which enables differentiable texture optimization through dual-renderer alignment while preserving the original simulation environment. To further ensure that the attack remains effective across long-horizon and diverse viewpoints in the physical world, we propose Trajectory-Aware Adversarial Optimization (TAAO), which prioritizes behaviorally critical frames and stabilizes optimization with a vertex-based parameterization. Built on these designs, we present Tex3D, the first framework for end-to-end optimization of 3D adversarial textures directly within the VLA simulation environment. Experiments in both simulation and real-robot settings show that Tex3D significantly degrades VLA performance across multiple manipulation tasks, achieving task failure rates of up to 96.7\%. Our empirical results expose critical vulnerabilities of VLA systems to physically grounded 3D adversarial attacks and highlight the need for robustness-aware training.

对抗攻击机器人安全3D纹理VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。