arXiv:2509.21657cs.CV2025-09被引 30

让视频模型生成有几何一致性的3D世界,无需逐场景调优。

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

  • 用可训练的几何分支增强冻结视频模型,统一建模视频与3D隐空间。
  • 在多视角一致性与风格保持上优于现有基线,提升显著。
  • 生成的隐向量可直接用于新视角合成等下游3D任务,通用性强。

高质量3D世界模型对具身智能和通用人工智能至关重要,支撑AR/VR内容创作与机器人导航等应用。尽管当前视频基础模型具备强大想象先验,但缺乏显式3D定位能力,导致空间一致性差且难以支持下游3D推理任务。本文提出FantasyWorld,一种通过可训练几何分支增强冻结视频基础模型的框架,实现视频潜在表示与隐式3D场的单次前向传播联合建模。方法引入跨分支监督:几何线索引导视频生成,视频先验正则化3D预测,从而获得一致且可泛化的3D感知视频表示。值得注意的是,几何分支生成的潜在表示可直接用于新视角合成、导航等下游3D任务,无需每场景优化或微调。大量实验表明,FantasyWorld有效弥合视频想象与3D感知的鸿沟,在多视角一致性与风格一致性上超越近期几何一致基线。消融研究进一步证实性能提升源于统一骨干网络与跨分支信息交换。

原文摘要 · Abstract (English)

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative priors, current video foundation models lack explicit 3D grounding capabilities, thus being limited in both spatial consistency and their utility for downstream 3D reasoning tasks. In this work, we present FantasyWorld, a geometry-enhanced framework that augments frozen video foundation models with a trainable geometric branch, enabling joint modeling of video latents and an implicit 3D field in a single forward pass. Our approach introduces cross-branch supervision, where geometry cues guide video generation and video priors regularize 3D prediction, thus yielding consistent and generalizable 3D-aware video representations. Notably, the resulting latents from the geometric branch can potentially serve as versatile representations for downstream 3D tasks such as novel view synthesis and navigation, without requiring per-scene optimization or fine-tuning. Extensive experiments show that FantasyWorld effectively bridges video imagination and 3D perception, outperforming recent geometry-consistent baselines in multi-view coherence and style consistency. Ablation studies further confirm that these gains stem from the unified backbone and cross-branch information exchange.

3D生成视频理解几何一致性多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。