通过操控想象诱导世界模型代理执行恶意行为,攻击隐蔽且持久。
TrojanWorld: Backdooring World-Model Agents via Imagination Steering

- 用物理物体作触发器,通过想象引导实现攻击
- 目标动作偏差低至0.026,干净性能保留超98.8%
- 攻击后移除触发器仍能持续执行恶意行为,适合研究安全威胁
世界模型日益成为基于模型强化学习代理的预测核心,使其能在行动前模拟未来动态并推理想象轨迹。其高昂的训练成本使预训练世界模型易于分发复用,暴露于供应链攻击风险中。后门攻击为利用此类供应链提供了针对性且隐蔽的途径,但对交互式世界模型代理的威胁尚未被充分探索。为此,我们提出TrojanWorld,一种针对世界模型代理的后门框架,通过操纵内部想象诱导攻击者指定行为。场景中放置的物理物体作为触发器,可通过代理原生观测流程在部署时激活,无需数字篡改观测流。为实现有效、隐蔽且持久的控制,TrojanWorld结合决策反馈引导的触发条件想象、清洁行为锚定以保持无触发下的预测与行为保真度,以及因果传播以维持触发消失后后续轨迹中的诱导偏好。这些机制共同构建了从物理感知到被污染想象再到恶意行为选择的端到端攻击链。在DeepMind Control、MetaWorld、MyoSuite和RoboDesk基准上,对TD-MPC2、DreamerV3和R2-Dreamer系统的实验表明,触发激活下,TrojanWorld实现最低0.026的目标动作偏差,同时保持至少98.8%的原始性能。即使触发器移除,受控代理仍可能持续执行攻击者指定动作。
原文摘要 · Abstract (English)
World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. Backdoor attacks offer a targeted and stealthy means of exploiting such supply chains, yet their threat to interactive world-model agents remains largely unexplored. To fill this gap, we present TrojanWorld, a backdoor framework for world-model agents that induces attacker-specified behavior by steering internal imagination. A physical object placed in the scene acts as the trigger, enabling deployment-time activation through the agent's native observation pipeline without digitally manipulating the observation stream. To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to sustain the induced preference along subsequent trajectories after the trigger disappears. Together, these mechanisms establish an end-to-end attack chain from physical perception through corrupted imagination to malicious action selection. Experiments with the TD-MPC2, DreamerV3, and R2-Dreamer systems across the DeepMind Control, MetaWorld, MyoSuite, and RoboDesk benchmarks show that under trigger activation, TrojanWorld achieves a target-action deviation as low as 0.026 while retaining at least 98.8% of the corresponding clean performance. Even after trigger removal, the compromised agent can remain trapped in the induced behavioral trajectory, continuing to execute attacker-specified actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。