arXiv:2608.16806cs.ROcs.AI2026-08

攻击者通过伪造环境状态文本,让大模型产生错误规划并执行恶意动作。

Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents

论文配图:Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents
图 1 · 摘自论文原文
  • 利用虚假环境状态信息诱导大模型生成错误计划。
  • 在多个场景中使任务成功率提升超89%,执行成功率提升43%。
  • 揭示了环境状态文本的欺骗性,适合安全与防御研究者参考。

基于大语言模型(LLM)的具身智能体依赖环境状态来理解场景、生成高层规划并驱动物理执行,因此规划可见的状态表征构成关键安全边界。现有攻击主要针对用户指令、提示上下文、模型行为或感知输入,却较少关注环境状态文本本身是否可作为误导性任务证据,并在规划后持续影响执行结果。由于具身任务受实体定位、动作前提、空间关系和环境约束限制,仅偏离规划不足以导致恶意执行。为此,我们首次将环境状态文本视为独立攻击面,提出闭环的环境状态-文本注入(ESTI)攻击。不修改原始指令、模型参数或执行器,ESTI将对抗目标重构为与当前环境兼容的虚假状态证据,通过对象属性、空间关系、可用性、任务阶段规则及执行反馈影响规划与执行。我们进一步构建ESTI-Bench评估攻击在规划到执行闭环中的传播能力,并在ProgPrompt/VirtualHome、VoxPoser/RLBench、AI2-THOR/iTHOR上对比ESTI与Vanilla IPI、EIRAD、BADROBOT。结果表明,ESTI持续优于基线,在规划级和执行级攻击成功率上分别提升最多达89.32%和43.69%。深入分析显示,定位一致性、逻辑一致性和可执行性共同决定被操纵的状态证据能否在具身闭环中传播并引发可验证的环境变化。

原文摘要 · Abstract (English)

Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive physical execution, making planner-visible state representations a critical security boundary. Existing attacks primarily manipulate user instructions, prompt contexts, model behavior, or perceptual inputs, while paying limited attention to whether environment-state text itself can serve as deceptive task evidence and propagate beyond planning to affect execution outcomes. Because embodied tasks are constrained by entity grounding, action preconditions, spatial relations, and environmental constraints, planning deviation alone does not guarantee adversarial execution. To address this gap, we investigate environment-state text as an independent attack surface and present the first closed-loop Environment State-Text Injection (ESTI) attack for LLM-driven embodied agents. Without modifying the original user instruction, model parameters, or executor, ESTI reformulates an adversarial objective as false state evidence compatible with the current environment and influences planning and execution through object properties, spatial relations, affordances, task-stage rules, and execution feedback. We further develop ESTI-Bench to evaluate attack propagation across the planning-to-execution closed loop and compare ESTI with Vanilla IPI, EIRAD, and BADROBOT across ProgPrompt/VirtualHome, VoxPoser/RLBench, and AI2-THOR/iTHOR. ESTI consistently outperforms existing baselines, improving planning-level and execution-level attack success rates by up to 89.32\% and 43.69\%, respectively. Further analysis shows that grounding, consistency, and executability jointly determine whether manipulated state evidence can propagate through the embodied closed loop and produce verifiable environmental changes.

具身智能安全攻击大模型环境欺骗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。