arXiv:2601.01075cs.LGcs.AI2026-01中稿 · ICML被引 7

让世界模型学会追踪动态环境中的流动对称性,提升部分观测下的长期预测能力。

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments

  • 利用时间参数化的对称性构建可迁移的潜在记忆,随自身运动和外部物体运动同步变化
  • 在2D/3D部分观测视频建模任务中,长期预测误差显著低于扩散、记忆增强和循环模型
  • 适合需要长期推理的机器人感知与具身智能系统,尤其在观测不全场景下表现优异

具身系统感知世界如同一曲‘流动的交响乐’:多种连续感官输入与自身运动交织,并与外部物体动态相互作用。这些感官流和世界底层动力学遵循平滑的时间参数对称性,但现有世界模型忽略此结构。缺乏尊重该结构的记忆机制,导致部分可观测性成为主要障碍——每次观测仅揭示世界片段,而未观测区域仍在持续演化。本文提出流动等变世界建模(Flow Equivariant World Modeling)框架,利用潜在记忆中的时间参数对称性,实现长时间跨度下的稳定准确动力学预测。该潜在记忆随自身运动及推断出的外部物体运动进行等变位移与变换,使视域外区域的信息在时间推进中保持对齐。我们在2D和3D部分观测视频建模基准上验证了该框架优于当前最先进的扩散模型、记忆增强模型与循环世界模型。更广泛地,结果表明,当预测表征与所建模世界的时空与动力学结构一致时,其表达能力更强。

原文摘要 · Abstract (English)

Embodied systems experience the world as 'a symphony of flows': a combination of many continuous streams of sensory input coupled to self-motion, interwoven with the dynamics of external objects. These sensory streams and the underlying dynamics of the world obey smooth, time-parameterized symmetries which existing world models ignore. Without a memory that respects this structure, partial observability presents a major obstacle to existing methods: each observation reveals only a fraction of the world, while unobserved regions continue to evolve. In this work, we introduce Flow Equivariant World Modeling, a framework that leverages time-parameterized symmetries within a latent memory for stable and accurate dynamics prediction over long horizons. The latent memory shifts and transforms equivariantly with self-motion and inferred external object motion, keeping information about out-of-view regions aligned as time progresses. We demonstrate the advantage of this framework over state-of-the-art diffusion, memory-augmented, and recurrent world model architectures on 2D and 3D partially observed video world modeling benchmarks. More broadly, our results suggest that predictive representations become more powerful when they are organized in line with the temporal and dynamical structure of the world they model. Project page: https://flowequivariantworldmodels.github.io/

世界模型动态建模具身智能记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。