arXiv:2602.18639cs.LGmath.OC2026-02被引 5

让视觉模型更抗干扰,提升规划稳定性。

Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models

  • 用双状态编码器强化控制相关状态等价性
  • 在背景变化下仍保持90%以上规划成功率
  • 适配多种预训练视觉特征,适合机器人导航

从高维视觉观测中学习的世界模型可使智能体直接在隐空间中决策与规划,避免像素级重建。然而,近期的隐式预测架构(如JEPAs)包括DINO世界模型(DINO-WM),在测试时因对“慢特征”敏感而出现鲁棒性下降,这些特征如背景变化、无关干扰物等与任务无关。本文通过在预测目标中引入双状态编码器,强制控制相关状态等价性,将具有相似转移动态的状态映射至相近的隐状态,同时抑制慢特征的影响。我们在不同背景变化和视觉干扰下的简单导航任务上评估该模型,结果表明,在所有基准测试中,该模型均显著提升对慢特征的鲁棒性,且隐空间大小缩小至DINO-WM的1/10。此外,该模型对预训练视觉编码器选择不敏感,与DINOv2、SimDINOv2和iBOT特征搭配时仍保持鲁棒性。

原文摘要 · Abstract (English)

World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to "slow features". These include visual variations such as background changes and distractors that are irrelevant to the task being solved. We address this limitation by augmenting the predictive objective with a bisimulation encoder that enforces control-relevant state equivalence, mapping states with similar transition dynamics to nearby latent states while limiting contributions from slow features. We evaluate our model on a simple navigation task under different test-time background changes and visual distractors. Across all benchmarks, our model consistently improves robustness to slow features while operating in a reduced latent space, up to 10x smaller than that of DINO-WM. Moreover, our model is agnostic to the choice of pretrained visual encoder and maintains robustness when paired with DINOv2, SimDINOv2, and iBOT features.

世界模型视觉规划不变表征鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。