5秒生成高质量3D模型,速度提升10倍,适合快速建模与训练。
Efficient 3D Content Reconstruction and Generation

- 融合多视角扩散与前馈稀疏重建,实现快速生成
- 5-20秒完成高精度3D资产生成,速度比现有方法快10倍
- 适用于游戏、虚拟现实等需快速生成3D内容的场景
自动3D内容生成旨在用文本或图像直接合成或恢复3D资产,替代耗时的手动建模和扫描流程。其应用涵盖视频游戏、虚拟现实、机器人和仿真,支持快速原型设计、多样化交互世界生成以及训练基础模型所需的高效3D数据采集。当前方法主要分为两类:(i) 文本或图像到3D生成,通过学习3D几何与外观先验,从自然语言或单张图像生成新资产;(ii) 3D重建,从RGB图像估计相机位姿与几何结构。本文在两个方向均取得进展:在生成方面,提出Instant3D,结合多视角扩散模型与前馈稀疏视图3D重建,在5-20秒内生成高质量资产;在重建方面,开发FastMap,采用一阶优化与融合GPU内核的结构-运动算法,相比最先进方法提速达10倍,同时保持相当的位姿精度与下游新视角合成质量。
原文摘要 · Abstract (English)
Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets directly from text or images. Its applications span video games, virtual reality, robotics, and simulation, enabling rapid asset prototyping, diverse interactive world generation, and efficient 3D data collection for training foundation models. Contemporary solutions largely follow two complementary paradigms: (i) text- or image-to-3D generation, which learns priors over 3D geometry and appearance to create novel assets from natural language or a single view image; and (ii) 3D reconstruction, which estimates camera poses and geometry from RGB images. This thesis advances both directions. On the generation side, I introduce Instant3D, which combines multi-view diffusion with feed-forward sparse-view 3D reconstruction to produce high-quality assets in 5-20 seconds. On the reconstruction side, I develop FastMap, a structure-from-motion pipeline that achieves up to 10x speedup over prior state-of-the-art by using first-order optimization with fused GPU kernels extensively, while maintaining comparable pose accuracy and downstream novel view synthesis quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。