在高维特征空间中实现流畅的随机世界建模,提升感知与多样性。
Flow Matching in Feature Space for Stochastic World Modeling

- 直接在预训练特征空间(如DINOv3)做流匹配,避免低维重构损失信息。
- 提出可微分一步投影机制,实现高效训练与时间一致性。
- 在真实与合成数据集上均提升感知性能、模式覆盖和长时鲁棒性。
世界建模需在预测不确定未来的同时保留对下游感知有用的信息。现有视觉世界模型常难以兼顾两者:基于变分自编码器的随机模型在低维重建隐空间运行,可能损害感知性能;而使用强预训练特征的确定性预测器则将多模态未来压缩为单一模糊均值。本文提出FlowWM,一种在预训练特征空间(如DINOv3)中直接进行流匹配的随机世界模型。该方法面临挑战,因预训练特征维度极高,标准扩散方案效果不佳。为此,我们研究了特征空间流匹配的设计选择,并引入可微分的一步投影机制,实现高效训练并保持时间一致性与任务驱动目标。我们在两个基准上评估:一个合成基准用于系统评估准确率与多样性,一个真实世界基准FuturePerception。结果表明,FlowWM在感知性能、模式覆盖和时序鲁棒性上均有提升,验证了高维特征空间中随机世界建模的有效设计。
原文摘要 · Abstract (English)
World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world models often struggle to satisfy both goals: VAE-based stochastic models operate in low-dimensional reconstruction latents, which can limit perception performance, while deterministic predictors using strong pretrained features collapse multimodal futures into a single blurry mean. In this work, we propose FlowWM, a stochastic world model that performs flow matching directly within pretrained feature space (e.g., DINOv3). This is challenging because pretrained features are substantially high-dimensional, making standard diffusion recipes suboptimal. To address this, we investigate the design choices needed for feature-space flow matching and introduce a differentiable one-step projection mechanism that enables efficient training with temporal consistency and task-driven objectives. We evaluate FlowWM on two benchmarks: a synthetic benchmark for systematic evaluation of accuracy and diversity, and a real-world benchmark FuturePerception. FlowWM improves perception performance, mode coverage, and horizon robustness, validating our proposed design for stochastic world modeling in high-dimensional feature spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。