arXiv:2507.19468cs.CV2025-07被引 56

用DINOv2的特征空间构建视频世界模型,预测未来帧并支持规划。

Back to the Features: DINO as a Foundation for Video World Models

  • 基于预训练DINOv2编码器,在大规模未标注视频上训练未来帧预测器。
  • 在分割、深度预测等任务上优于现有模型,具备直观物理理解能力。
  • 可微调为动作条件模型,用于潜在空间中的轨迹模拟与规划。

我们提出DINO-world,一种强大的通用视频世界模型,通过DINOv2的潜在空间预测未来帧。利用预训练图像编码器,并在大规模未标注视频数据集上训练未来预测器,DINO-world学习了从驾驶、室内场景到仿真环境等多种场景的时序动态。实验表明,DINO-world在多个视频预测基准上表现优于先前模型,例如分割和深度预测任务,并展现出对直观物理的强理解能力。此外,我们证明可在观测-动作轨迹上微调预测器,得到的动作条件世界模型可用于规划——在潜在空间中模拟候选轨迹。

原文摘要 · Abstract (English)

We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and training a future predictor on a large-scale uncurated video dataset, DINO-world learns the temporal dynamics of diverse scenes, from driving and indoor scenes to simulated environments. We show that DINO-world outperforms previous models on a variety of video prediction benchmarks, e.g. segmentation and depth forecasting, and demonstrates strong understanding of intuitive physics. Furthermore, we show that it is possible to fine-tune the predictor on observation-action trajectories. The resulting action-conditioned world model can be used for planning by simulating candidate trajectories in latent space.

视频预测世界模型DINOv2动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。