提出可执行的越狱攻击框架,揭示大模型机器人安全漏洞。
POEX: Towards Policy Executable Jailbreak Attacks Against the LLM-based Robots
- 设计POEX框架,通过梯度优化生成可执行的有害策略
- 实测验证攻击在真实机器人与仿真环境中的有效性
- 适合关注机器人安全、大模型防御的研究者
大语言模型(LLM)在机器人中的集成日益广泛,可将指令转化为可执行的机器人策略。然而,LLM易受越狱攻击,可能引发物理世界中的安全风险。本文探究了针对基于LLM的机器人实施越狱攻击的可行性与机制,回答三个问题:(1) 现有LLM越狱攻击在机器人场景中是否适用?(2) 若不直接适用,会面临哪些独特挑战?(3) 如何防御此类攻击?为此,我们构建了面向人类-物体-环境风险的有害RLBench数据集,并对基于LLM的机器人系统进行了测量研究。结果表明,传统越狱攻击在机器人场景中不可行,主要面临两大挑战:确定可执行策略的优化方向,以及准确评估策略的实际可执行性。为实现更全面的安全分析,我们提出POEX(Policy Executable)越狱框架,通过隐藏层梯度优化确保越狱成功与策略可执行,并引入多智能体评估器精准衡量策略可行性。在真实机器人系统和仿真环境中进行的实验验证了POEX的有效性,揭示了关键安全漏洞及其在不同LLM间的迁移能力。最后,我们提出基于提示和基于模型的防御方案。研究强调了在关键应用中保障基于LLM机器人的安全性刻不容缓。
原文摘要 · Abstract (English)
The integration of LLMs into robots has witnessed significant growth, where LLMs can convert instructions into executable robot policies. However, the inherent vulnerability of LLMs to jailbreak attacks brings critical security risks from the digital domain to the physical world. An attacked LLM-based robot could execute harmful policies and cause physical harm. In this paper, we investigate the feasibility and rationale of jailbreak attacks against LLM-based robots and answer three research questions: (1) How applicable are existing LLM jailbreak attacks against LLM-based robots? (2) What unique challenges arise if they are not directly applicable? (3) How to defend against such jailbreak attacks? To this end, we first construct a "human-object-environment" robot risks-oriented Harmful-RLbench and then conduct a measurement study on LLM-based robot systems. Our findings conclude that traditional LLM jailbreak attacks are inapplicable in robot scenarios, and we identify two unique challenges: determining policy-executable optimization directions and accurately evaluating robot-jailbroken policies. To enable a more thorough security analysis, we introduce POEX (POlicy EXecutable) jailbreak, a red-teaming framework that induces harmful yet executable policy to jailbreak LLM-based robots. POEX incorporates hidden layer gradient optimization to guarantee jailbreak success and policy execution as well as a multi-agent evaluator to accurately assess the practical executability of policies. Experiments conducted on the real-world robotic systems and in simulation demonstrate the efficacy of POEX, highlighting critical security vulnerabilities and its transferability across LLMs. Finally, we propose prompt-based and model-based defenses to mitigate attacks. Our findings underscore the urgent need for security measures to ensure the safe deployment of LLM-based robots in critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。