用DINOv2的特征空间构建视频世界模型,预测未来帧并支持规划。
Back to the Features: DINO as a Foundation for Video World Models
- 基于预训练DINOv2编码器,在大规模未标注视频上训练未来帧预测器。
- 在分割、深度预测等任务上优于现有模型,具备直观物理理解能力。
- 可微调为动作条件模型,用于潜在空间中的轨迹模拟与规划。
我们提出DINO-world,一种强大的通用视频世界模型,通过DINOv2的潜在空间预测未来帧。利用预训练图像编码器,并在大规模未标注视频数据集上训练未来预测器,DINO-world学习了从驾驶、室内场景到仿真环境等多种场景的时序动态。实验表明,DINO-world在多个视频预测基准上表现优于先前模型,例如分割和深度预测任务,并展现出对直观物理的强理解能力。此外,我们证明可在观测-动作轨迹上微调预测器,得到的动作条件世界模型可用于规划——在潜在空间中模拟候选轨迹。
原文摘要 · Abstract (English)
We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and training a future predictor on a large-scale uncurated video dataset, DINO-world learns the temporal dynamics of diverse scenes, from driving and indoor scenes to simulated environments. We show that DINO-world outperforms previous models on a variety of video prediction benchmarks, e.g. segmentation and depth forecasting, and demonstrates strong understanding of intuitive physics. Furthermore, we show that it is possible to fine-tune the predictor on observation-action trajectories. The resulting action-conditioned world model can be used for planning by simulating candidate trajectories in latent space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。