arXiv:2504.17070cs.ROcs.AI2025-04中稿 · publication at the…被引 5

提出多触发后门攻击,可让机器人在特定指令下执行恶意行为

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

  • 通过少量任务参数注入后门,隐蔽控制机器人决策
  • 设计多触发词优化方法,适配不同机器人应用场景
  • 首次针对大模型驱动的机器人规划系统构建攻击框架

机器人需依赖任务规划实现多步操作。近期大型预训练模型在任务规划中表现优异,如大语言模型(LLMs)可基于动作与目标描述生成任务计划。尽管大模型在机器人智能方面进展迅速,其安全影响仍不明确,存在重要漏洞研究空白。本文提出MuTRAP,首个专为大模型辅助机器人任务规划系统设计的多触发后门攻击方法。该方法遵循机器人领域常用模式:骨干大模型通常被冻结并部署于中心服务器,限制攻击者直接访问。相比之下,MuTRAP通过少量任务特定参数注入后门。此外,我们开发了触发词优化方法,用于选择对不同机器人应用最有效的多触发词。例如,使用唯一触发词“herical”可激活特定恶意行为,如厨房机器人误切手。通过展示当前基于大模型规划系统的脆弱性,本研究旨在推动安全机器人智能的发展。更多细节与演示见:https://mutrap.github.io/MuTRAP/

原文摘要 · Abstract (English)

Robots need task planning methods to achieve goals that require more than one action. Recently, large pretrained models have demonstrated impressive performance in task planning. For instance, large language models (LLMs) can generate task plans using action and goal descriptions. Despite the rapid progress of large models in robot intelligence, their security implications remain only partially understood, leaving important gaps in the exploration of potential vulnerabilities in LLM-driven robotic planning systems. To investigate such risks, in this paper, we develop MuTRAP, the first multi-trigger trojan attack specifically designed and targeted for LLM-assisted robot task planners. MuTRAP follows the standard practice of LLM usage in robotics where the backbone LLM is typically frozen and hosted in a central server limiting attacker's reach. In contrast, MuTRAP injects backdoor using a small set of task-specific parameters. In addition, we develop a trigger optimization method for selecting multiple-trigger words that are most effective for different robot applications. For instance, one can use unique trigger word "herical" to activate a specific malicious behavior, e.g., cutting hand on a kitchen robot. Through MuTRAP that demonstrates the vulnerability of current LLM-based planners, our goal is to promote the development of secured robot intelligence. Details and demos are provided in: https://mutrap.github.io/MuTRAP/

机器人安全后门攻击大模型任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。