arXiv:2606.12978cs.ROcs.CV2026-06

看似正常的指令可能暗中操控机器人最终动作,暴露语言控制漏洞。

Trajectory-Level Redirection Attacks on Vision-Language-Action Models

论文配图:Trajectory-Level Redirection Attacks on Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用闭环推理搜索近似原指令的干扰提示
  • 仅改写少量词语即可让机器人完成攻击者指定任务
  • 硬件实验验证了对真实机器人的隐蔽操控能力

视觉-语言-动作(VLA)策略将自然语言引入闭环机器人控制,使机器人可直接根据文本指令执行操作任务。同一接口也使文本在控制中反复出现,因为提示在每次重规划步骤中被重复使用,且每个提示驱动的动作会改变策略后续观察到的环境状态。现有VLA攻击研究关注能诱发特定低级动作或使动作持续存在的对抗性提示。我们识别出更强的轨迹级失效模式:一个看似仍指定正确任务但实际引导最终物理结果偏离的提示。我们将其数学形式化为“命令保持型轨迹重定向”,一种仅通过提示攻击的威胁模型,攻击者在对话开始前选择一个提示,所有策略与环境组件保持不变,且提示必须与原始指令接近,不能包含目标词或修正语言。为发现此类提示,我们提出一种基于策略的提示搜索方法,利用轨迹回放发现其闭环行为符合目标任务且满足命令保持约束的扰动。仿真和硬件实验表明,接近原指令的提示扰动可成功将VLA轨迹重定向至攻击者指定的目标。这些结果揭示了VLA指令对齐中的轨迹级漏洞:看似保留原始指令的文本仍可能让攻击者掌控机器人最终的物理输出。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies bring natural language into closed-loop robot control, enabling robots to execute manipulation tasks directly from text instructions. The same interface gives text a recurring role in control because the prompt is reused at every replanning step, and each prompt-conditioned action changes the future observations on which the policy acts. Existing VLA attacks study adversarial prompts that elicit targeted low-level actions or make such actions persist across changing images. We identify a stronger trajectory-level failure mode: a prompt that still $\textit{appears}$ to specify the intended task but redirects the final physical outcome. We mathematically formalize this setting as $\textit{command-preserving trajectory redirection}$, a prompt-only threat model in which the attacker chooses one prompt before the episode, all policy and environment components remain fixed, and the prompt must stay close to the benign instruction while omitting target words and correction language. To find such prompts, we introduce an on-policy prompt search method that uses rollouts to discover perturbations whose closed-loop behavior tracks a target task while satisfying the command-preserving constraints. Experiments in simulation and on hardware show that near-benign prompt perturbations can redirect VLA rollouts to attacker-specified targets. These results expose a trajectory-level vulnerability in VLA instruction grounding: text that appears to preserve the intended command can still give an adversary control over the robot's final physical outcome. Project website: https://vla-redirection-attack.github.io/

机器人安全指令操控对抗攻击VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。