用一张纹理图让多任务机器人模型误操作,攻击效果跨任务通用。
UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models

- 设计统一纹理,通过可微渲染反向优化跨任务攻击
- 攻击使任务成功率从90.0%降至48.4%,动作偏离目标
- 无需重调,攻击可跨模型、跨任务迁移,适合安全评估
视觉-语言-动作(VLA)模型作为通用机器人策略,能响应多样语言指令并完成多种操作任务。然而其对实体代理的直接控制也使其易受对抗干扰,可能导致不安全行为。现有攻击通常针对单一任务,未充分探索多任务VLA的跨任务脆弱性。本文提出UniTexture,一种基于单个3D纹理物体的跨任务通用对抗纹理攻击方法,通过可微渲染将策略输出梯度反传至表面纹理参数,联合优化共享纹理以覆盖多种任务、指令、状态和视角,采用目标动作空间损失函数,引导预测动作朝攻击者设定目标偏移,无需为每项任务单独优化。在OpenVLA和$π_{0.5}$上评估显示,该攻击将平均任务成功率从90.0%降至48.4%,引发目标对齐的动作偏移,并展现出无需重优化即可实现跨套件与跨模型迁移的能力。这些结果揭示了多任务VLA存在可被单一对抗表面纹理系统性利用的共性漏洞。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe physical behaviors. Existing attacks on robotic policies are typically optimized for a single task or instruction, leaving the cross-task vulnerabilities of multitask VLAs largely unexplored. We introduce UniTexture, a cross-task universal adversarial texture attack that uses a single textured 3D object to induce targeted deviations in VLA action predictions across multiple tasks. UniTexture backpropagates gradients from the policy's action outputs to surface texture parameters through a differentiable renderer. It jointly optimizes the shared texture over a distribution of tasks, instructions, states, and viewpoints using a targeted action-space objective, steering predicted actions toward attacker-defined targets without optimizing a separate texture for each task. We evaluate UniTexture on OpenVLA and $π_{0.5}$ across diverse manipulation tasks and multiple evaluation settings. UniTexture reduces the mean task success rate from 90.0% under benign conditions to 48.4% under attack, induces target-aligned action shifts, and further exhibits cross-suite and cross-model transfer without re-optimization. Together, these findings reveal shared cross-task vulnerabilities in multitask VLAs that can be systematically exploited through a single adversarial surface texture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。