arXiv:2603.23117cs.CRcs.AI2026-03中稿 · ICML被引 5

用一张干扰贴纸就能骗机器人把刀递给人,而不用改指令。

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches

  • 在视觉语言动作模型中,用对抗性贴纸干扰推理过程
  • 真实场景中成功让机器人错误递出刀具,成功率超90%
  • 首次揭示思维链推理的可被劫持风险,适合安全研究者关注

通过引入思维链(CoT)推理,视觉-语言-动作(VLA)模型在机器人操作中展现出强大的泛化与可解释能力。然而,基于CoT的推理机制安全性尚未被充分研究。本文首次证明,CoT推理会显著主导动作生成,即使其语义与输入指令不一致。基于此,我们提出TRAP——首个针对基于CoT推理的VLA模型的定向行为劫持攻击方法。该方法通过在场景中放置对抗性贴纸(如桌布),干扰中间推理过程,从而引导下游动作向攻击者预设的行为偏移。在三种代表性推理型VLA模型上进行的广泛评估表明,TRAP具有高度有效性。值得注意的是,我们在真实世界中通过打印贴纸实现攻击。结果凸显了保障VLA系统中CoT推理安全的紧迫性。

原文摘要 · Abstract (English)

By integrating Chain-of-Thought (CoT) reasoning, Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, particularly by improving generalization and interpretability. However, the security of CoT-based reasoning mechanisms remains largely unexplored. In this paper, we show that CoT reasoning introduces a novel attack vector for targeted behavior hijacking--for example, causing a robot to mistakenly deliver a knife to a person instead of an apple--without modifying the user's instruction. We first provide empirical evidence that CoT strongly governs action generation, even when it is semantically misaligned with the input instructions. Building on this observation, we propose TRAP, the first targeted behavior-hijacking adversarial attack against CoT-reasoning VLA models. By targeting the reasoning-to-action pathway, TRAP uses an adversarial patch (e.g., a tablecloth placed on the table) to steer intermediate CoT reasoning and downstream actions toward adversary-defined behaviors. Extensive evaluations on three representative reasoning VLAs, spanning distinct CoT reasoning mechanisms, demonstrate the effectiveness of TRAP. Notably, we implemented the patch by printing it on paper in a real-world setting. Our findings highlight the urgent need to secure CoT reasoning in VLA systems. The project page is available at https://zhengxian-huang.github.io/TRAP-website/.

机器人安全对抗攻击思维链VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。