arXiv:2603.01414cs.RO2026-03中稿 · ACM SenSys 2026被引 3

通过动作级操纵,让机器人执行看似安全实则危险的指令。

Jailbreaking Embodied LLMs via Action-level Manipulation

  • 用代理模型生成表面安全的动作序列,欺骗机器人执行危险操作。
  • 在仿真和真实机械臂上实现53%更高的攻击成功率。
  • 适合研究机器人安全与对抗性攻击的学者参考。

具身大语言模型(LLMs)使智能体能通过自然语言指令和动作与物理世界交互。然而,除了大语言模型固有的语言层面风险外,具备真实动作执行能力的具身LLM引入了新漏洞:看似语义安全的指令仍可能导致危险的物理后果,暴露出语言安全与物理结果之间的根本错位。本文提出Blindfold,一种自动化攻击框架,利用具身LLM在真实动作情境中有限的因果推理能力。不同于对黑箱具身LLM进行迭代试错式越狱,Blindfold采用对抗性代理规划策略:通过本地替代模型执行看似安全但可能造成危害的动作操纵。Blindfold进一步通过注入精心设计的噪声隐藏恶意动作,以规避防御机制,并集成规则验证器提升攻击可执行性。在具身AI模拟器和真实6自由度机械臂上的评估表明,Blindfold相比现有最先进基线攻击成功率最高提升53%,凸显了必须超越表层语言过滤,转向后果感知型防御机制以保障具身LLM安全的紧迫性。

原文摘要 · Abstract (English)

Embodied Large Language Models (LLMs) enable AI agents to interact with the physical world through natural language instructions and actions. However, beyond the language-level risks inherent to LLMs themselves, embodied LLMs with real-world actuation introduce a new vulnerability: instructions that appear semantically benign may still lead to dangerous real-world consequences, revealing a fundamental misalignment between linguistic security and physical outcomes. In this paper, we introduce Blindfold, an automated attack framework that leverages the limited causal reasoning capabilities of embodied LLMs in real-world action contexts. Rather than iterative trial-and-error jailbreaking of black-box embodied LLMs, Blindfold adopts an Adversarial Proxy Planning strategy: it compromises a local surrogate LLM to perform action-level manipulations that appear semantically safe but could result in harmful physical effects when executed. Blindfold further conceals key malicious actions by injecting carefully crafted noise to evade detection by defense mechanisms, and it incorporates a rule-based verifier to improve the attack executability. Evaluations on both embodied AI simulators and a real-world 6DoF robotic arm show that Blindfold achieves up to 53% higher attack success rates than SOTA baselines, highlighting the urgent need to move beyond surface-level language censorship and toward consequence-aware defense mechanisms to secure embodied LLMs.

具身智能对抗攻击机器人安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。