arXiv:2605.01950cs.LGcs.AI2026-05被引 3

针对世界模型的长尾轨迹排序漏洞,提出新型后门攻击框架TRAP。

TRAP: Tail-aware Ranking Attack for World-Model Planning

论文配图:TRAP: Tail-aware Ranking Attack for World-Model Planning
图 1 · 摘自论文原文
  • 通过关注关键轨迹的排序优化,精准扰动决策路径。
  • 在多任务下使智能体持续偏离正常行为,性能显著下降。
  • 适用于评估基于世界模型智能体的安全性,尤其关注长期规划场景。

世界模型通过内部生成并评估想象中的轨迹,实现长时程规划,是通用智能体的重要基础。然而,这种以想象驱动的决策过程也引入了新的安全风险。现有后门攻击通常针对局部特征、单步预测或即时策略输出,对世界模型效果有限,因其动态先验和规划过程能吸收浅层扰动。我们发现,世界模型存在独特漏洞:想象轨迹的长尾排序结构中,仅扰动少数关键轨迹的顺序即可系统性劫持规划结果。为此,我们提出TRAP框架,通过尾部感知排序损失聚焦于关键轨迹,并结合双门控机制稳定优化与控制攻击施加时机和位置。触发条件下,TRAP改变想象轨迹的相对排序以引导规划偏离,同时保持干净输入下的正常排序结构。在DreamerV3和TD-MPC2上跨多种任务的实验表明,TRAP持续引发行为偏差并导致显著性能下降,凸显对基于世界模型的智能体进行专项安全评估的必要性。

原文摘要 · Abstract (English)

World models enable long-horizon planning by internally generating and evaluating imagined trajectories, making them a promising foundation for generalist agents. However, this imagination-driven decision process also introduces new security risks. Existing backdoor attacks typically aim to manipulate local features, one-step predictions, or instantaneous policy outputs. While such objectives may suffice for weaker reactive models, they are often ineffective against world models, where the learned dynamics prior and planning process can absorb or wash out the effects of shallow perturbations. More importantly, we find that world models exhibit a distinct backdoor vulnerability rooted in the long-tailed ranking structure of imagined trajectories, where disrupting the ordering of a few decision-critical trajectories can systematically hijack planning. To exploit this vulnerability, we propose TRAP, a backdoor attack framework for world models that targets imagined trajectory ranking. TRAP combines a tail-aware ranking loss to focus optimization on decision-critical trajectories with dual gating mechanisms that stabilize optimization and regulate when and where the attack penalty is applied. Under trigger conditions, TRAP alters the relative ranking of imagined trajectories to redirect planning outcomes, while largely maintaining the normal ranking structure on clean inputs. Experiments on DreamerV3 and TD-MPC2 across diverse tasks show that TRAP consistently induces sustained behavioral deviations and significant performance degradation, highlighting the need for dedicated security evaluation of world-model-based agents.

世界模型后门攻击规划安全长尾排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。