arXiv:2502.06608cs.CVcs.AI2025-02TPAMI被引 204

用大规模流模型生成高保真3D形状,对齐输入图像更精准。

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

论文配图:TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
图 1 · 摘自论文原文
  • 基于大规模校正流变换器,提升3D生成质量。
  • 200万高质量3D数据训练,实现高精度重建。
  • 适合需要逼真3D建模的AI应用开发者使用。

近期扩散技术在图像与视频生成上取得显著进展,大幅推动生成式AI的应用落地。然而,3D形状生成仍受制于数据规模有限、处理复杂及领域内先进方法探索不足。现有方法在输出质量、泛化能力与输入对齐方面存在明显短板。本文提出TripoSG,一种新型高效3D形状扩散范式,可生成与输入图像精确对应的高保真3D网格。具体贡献包括:1)构建大规模校正流变换器用于3D生成,在海量高质量数据上训练,实现当前最优保真度;2)采用融合SDF、法线与eikonal损失的混合监督策略训练3D VAE,提升重建质量;3)设计数据处理流程,生成200万高质量3D样本,揭示数据量与质量对训练3D生成模型的关键作用。通过全面实验验证各组件有效性,整体框架达到3D生成领域最新水平。生成的3D形状细节丰富,高分辨率下仍保持高保真,且对不同图像风格与内容具备优异泛化能力。为促进3D生成领域发展,模型将公开发布。

原文摘要 · Abstract (English)

Recent advancements in diffusion techniques have propelled image and video generation to unprecedented levels of quality, significantly accelerating the deployment and application of generative AI. However, 3D shape generation technology has so far lagged behind, constrained by limitations in 3D data scale, complexity of 3D data processing, and insufficient exploration of advanced techniques in the 3D domain. Current approaches to 3D shape generation face substantial challenges in terms of output quality, generalization capability, and alignment with input conditions. We present TripoSG, a new streamlined shape diffusion paradigm capable of generating high-fidelity 3D meshes with precise correspondence to input images. Specifically, we propose: 1) A large-scale rectified flow transformer for 3D shape generation, achieving state-of-the-art fidelity through training on extensive, high-quality data. 2) A hybrid supervised training strategy combining SDF, normal, and eikonal losses for 3D VAE, achieving high-quality 3D reconstruction performance. 3) A data processing pipeline to generate 2 million high-quality 3D samples, highlighting the crucial rules for data quality and quantity in training 3D generative models. Through comprehensive experiments, we have validated the effectiveness of each component in our new framework. The seamless integration of these parts has enabled TripoSG to achieve state-of-the-art performance in 3D shape generation. The resulting 3D shapes exhibit enhanced detail due to high-resolution capabilities and demonstrate exceptional fidelity to input images. Moreover, TripoSG demonstrates improved versatility in generating 3D models from diverse image styles and contents, showcasing strong generalization capabilities. To foster progress and innovation in the field of 3D generation, we will make our model publicly available.

3D生成扩散模型高保真形状合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。