arXiv:2508.08086cs.CVcs.GR2025-08被引 32

从单图或文本生成全景可探索3D世界,支持大范围、高保真场景重建。

Matrix-3D: Omnidirectional Explorable 3D World Generation

论文配图:Matrix-3D: Omnidirectional Explorable 3D World Generation
图 1 · 摘自论文原文
  • 用全景视频扩散模型结合场景网格条件,生成高质量连贯的3D视频。
  • 构建116K条带深度与轨迹标注的合成数据集,支持高效训练。
  • 提供快速重建与精确优化两种方案,适配不同精度需求的应用场景。

从单张图像或文本提示生成可探索的3D世界是空间智能的核心。现有方法虽利用视频模型实现广域通用生成,但常受限于生成场景范围。本文提出Matrix-3D框架,通过全景表征实现全覆盖的全景可探索3D世界生成,融合条件视频生成与全景3D重建。首先训练一种轨迹引导的全景视频扩散模型,以场景网格渲染为条件,实现高质量且几何一致的场景视频生成。为将全景视频升维至3D世界,提出两种方法:(1) 前馈式大全景重建模型,实现快速3D场景重建;(2) 基于优化的管线,实现高精度细节重建。为支持有效训练,还引入Matrix-Pano数据集,首个大规模合成数据集,包含116,000条高质量静态全景视频序列,附带深度与轨迹标注。大量实验表明,本框架在全景视频生成与3D世界生成任务上均达到当前最优性能。

原文摘要 · Abstract (English)

Explorable 3D world generation from a single image or text prompt forms a cornerstone of spatial intelligence. Recent works utilize video model to achieve wide-scope and generalizable 3D world generation. However, existing approaches often suffer from a limited scope in the generated scenes. In this work, we propose Matrix-3D, a framework that utilize panoramic representation for wide-coverage omnidirectional explorable 3D world generation that combines conditional video generation and panoramic 3D reconstruction. We first train a trajectory-guided panoramic video diffusion model that employs scene mesh renders as condition, to enable high-quality and geometrically consistent scene video generation. To lift the panorama scene video to 3D world, we propose two separate methods: (1) a feed-forward large panorama reconstruction model for rapid 3D scene reconstruction and (2) an optimization-based pipeline for accurate and detailed 3D scene reconstruction. To facilitate effective training, we also introduce the Matrix-Pano dataset, the first large-scale synthetic collection comprising 116K high-quality static panoramic video sequences with depth and trajectory annotations. Extensive experiments demonstrate that our proposed framework achieves state-of-the-art performance in panoramic video generation and 3D world generation. See more in https://matrix-3d.github.io.

3D生成全景重建视频扩散空间智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。