用误导性提示让机器人乱动,突破了传统攻击瓶颈
Rethinking the Intermediate Features in Adversarial Attacks: Misleading Robotic Models via Adversarial Distillation
- 用连续动作表示优化对抗前缀,避开离散动作限制
- 在13个任务上成功诱导机器人执行错误操作,成功率超基准方法
- 适合研究机器人安全与对抗攻击的学者参考
语言驱动的机器人学习通过单一模型响应语音指令显著提升了机器人适应性。然而,该领域安全性漏洞仍鲜有研究。本文提出一种针对语言条件机器人模型的新式对抗提示攻击,通过构造通用对抗前缀,使任何原始指令添加后均引发模型执行非预期动作。我们发现现有对抗技术在机器人场景中效果有限,主要因离散化动作空间具有天然鲁棒性。为此,提出基于连续动作表示优化对抗前缀,绕过离散化过程。此外,识别出中间特征对攻击的促进作用,利用中间自注意力特征的负梯度进一步提升攻击效力。在13个机器人操作任务上的实验验证了本方法在VIMA模型上的优越性及跨模型变体的迁移能力。
原文摘要 · Abstract (English)
Language-conditioned robotic learning has significantly enhanced robot adaptability by enabling a single model to execute diverse tasks in response to verbal commands. Despite these advancements, security vulnerabilities within this domain remain largely unexplored. This paper addresses this gap by proposing a novel adversarial prompt attack tailored to language-conditioned robotic models. Our approach involves crafting a universal adversarial prefix that induces the model to perform unintended actions when added to any original prompt. We demonstrate that existing adversarial techniques exhibit limited effectiveness when directly transferred to the robotic domain due to the inherent robustness of discretized robotic action spaces. To overcome this challenge, we propose to optimize adversarial prefixes based on continuous action representations, circumventing the discretization process. Additionally, we identify the beneficial impact of intermediate features on adversarial attacks and leverage the negative gradient of intermediate self-attention features to further enhance attack efficacy. Extensive experiments on VIMA models across 13 robot manipulation tasks validate the superiority of our method over existing approaches and demonstrate its transferability across different model variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。