研究大模型机器人系统中的提示注入攻击,发现攻击可导致错误动作并传播到多个机器人。
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
- 通过任务指令和感知模块间接注入恶意提示,触发异常行为。
- 攻击在多智能体间传播,任务完成率下降30%以上,具体取决于提示结构。
- 首次系统分析多智能体机器人中提示攻击的传播机制与防御弱点。
大型语言模型正越来越多地被集成到自主机器人系统中用于任务规划与控制,但这种集成使系统面临提示注入攻击,可能导致不安全决策甚至物理伤害。多智能体场景通过跨智能体污染和更广的攻击面增加了风险。本文评估了基于大模型的多智能体机器人系统所面临的提示注入攻击,既考虑对任务指令的直接注入,也考察通过感知模块的间接注入。在不同攻击目标复杂度和注入策略下,于单智能体与多智能体环境中进行实验,结果表明提示注入可诱导敌对行为并降低任务完成率。我们发现攻击可通过共享提示结构从一个智能体传播至其他智能体,影响程度取决于提示组成与目标智能体。此外,我们分析了架构变化如何影响大模型查询,进而影响攻击成功率。据我们所知,这是首个系统研究基于大模型的多智能体机器人系统中提示注入攻击的研究。
原文摘要 · Abstract (English)
Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to prompt injection attacks that can lead to unsafe decisions and physical harm. Multi-agent settings increase the risks through cross-agent contamination and broader attack surfaces. In this paper, we evaluate prompt injection attacks against an LLM-based multi-agent robotic system, considering both direct injections into task instructions and indirect injections through perception modules. In our experiments across varying attack-goal complexities and injection strategies in both single-agent and multi-agent settings, we show that prompt injection can induce adversarial actions while reducing task completion. We find that attacks can propagate from one agent to others through shared prompt structures, with impacts varying depending on prompt composition and the targeted agent. We further analyze how architectural changes affect LLM queries and, consequently, the attack success. To the best of our knowledge, this is the first study that systematically investigates prompt injection attacks in a multi-agent LLM-based robotic system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。