世界模型可被恶意注入,生成危险机器人训练数据。
Targeting World Models to Compromise Robot Learning Pipelines

- 在看似安全的操控数据中植入隐蔽恶意指令,仅在世界模型中触发
- 攻击成功生成合成危险轨迹,导致下游强化学习策略失效
- 适用于先进动作/文本条件世界模型,威胁机器人安全部署
世界模型近年来因其高效生成机器人训练数据或模拟真实环境的能力而迅速普及,被广泛集成到机器人学习流程中。然而,本文揭示其引入了一种独特且隐蔽的数据投毒入口,即使使用看似安全的真实数据训练,仍可能导致部署不安全或被破坏的机器人策略。与传统直接注入危险轨迹的方法不同,本研究提出的攻击方法将恶意提示或破坏性转换动态嵌入表面无害的遥操作数据中,仅在通过世界模型处理时才激活。这会生成合成的危险训练轨迹,进而导致下游强化学习策略产生后门行为。我们在最先进的动作条件和文本条件世界模型上验证了该攻击的有效性,实现了端到端的深度强化学习策略后门,并在视觉语言动作(VLA)设置中提供了概念验证。这些发现凸显了对更安全世界模型的研究必要性,并要求重新评估其在机器人学习供应链中的角色。
原文摘要 · Abstract (English)
World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed through a world model as input. This can result in the generation of synthetic, dangerous robot training trajectories and subsequently unsafe or compromised robot policies. We demonstrate the effectiveness of our attacks against both state of the art action conditioned and text conditioned world models, showing a full end-to-end backdoor on a downstream DRL policy and a proof-of-concept for the VLA setting. Overall these findings necessitate research into more secure world models and reevaluating their position within the robot learning supply chain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。