arXiv:2601.00393cs.CV2026-01被引 33

用单目视频实现可扩展的4D世界建模与生成

NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos

  • 无需姿态信息,直接重建4D场景,支持野外单目视频
  • 在线模拟单目退化模式,提升真实场景泛化能力
  • 性能领先现有方法,适合多领域应用开发

本文提出NeoVerse,一种可扩展的4D世界模型,支持4D重建、新轨迹视频生成及丰富下游应用。现有4D建模方法受限于昂贵的多视角4D数据或复杂的预处理流程,导致难以推广。NeoVerse基于无姿态前馈重建、在线单目退化模式模拟等设计,实现对多样化野外单目视频的高效适配。该架构兼具通用性与强泛化能力,在标准重建与生成基准上达到当前最优表现。项目主页见https://neoverse-4d.github.io。

原文摘要 · Abstract (English)

In this paper, we propose NeoVerse, a versatile 4D world model that is capable of 4D reconstruction, novel-trajectory video generation, and rich downstream applications. We first identify a common limitation of scalability in current 4D world modeling methods, caused either by expensive and specialized multi-view 4D data or by cumbersome training pre-processing. In contrast, our NeoVerse is built upon a core philosophy that makes the full pipeline scalable to diverse in-the-wild monocular videos. Specifically, NeoVerse features pose-free feed-forward 4D reconstruction, online monocular degradation pattern simulation, and other well-aligned techniques. These designs empower NeoVerse with versatility and generalization to various domains. Meanwhile, NeoVerse achieves state-of-the-art performance in standard reconstruction and generation benchmarks. Our project page is available at https://neoverse-4d.github.io.

4D建模单目视频视频生成泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。