arXiv:2501.16764cs.CV2025-01ICLR被引 62

用图像扩散模型生成高质量3D点云,解决多视角不一致问题

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

  • 直接利用大模型生成3D高斯点云,融合网页级2D先验知识
  • 通过多视角渲染损失确保任意视角下3D一致性,生成质量更高
  • 兼容现有图像生成技术,适合想快速迁移2D生成能力到3D的研究者

近期基于文本或单图的3D内容生成面临高质量3D数据集稀缺及2D多视角生成不一致的问题。本文提出DiffSplat,一种原生生成3D高斯点云的新框架,通过驯服大规模文生图扩散模型实现。与以往方法不同,该框架有效利用网络级2D先验知识,同时在统一模型中保持3D一致性。为启动训练,设计轻量级重建模型,可快速生成多视角高斯点云网格,支持大规模数据集构建。结合网格上的标准扩散损失与新增的3D渲染损失,促进任意视角下的3D一致性。该框架与图像扩散模型高度兼容,可无缝迁移多种图像生成技术至3D领域。大量实验表明,DiffSplat在文本和图像条件生成任务及下游应用中表现更优。详尽消融实验验证了各项设计的有效性,并揭示其内在机制。

原文摘要 · Abstract (English)

Recent advancements in 3D content generation from text or a single image struggle with limited high-quality 3D datasets and inconsistency from 2D multi-view generation. We introduce DiffSplat, a novel 3D generative framework that natively generates 3D Gaussian splats by taming large-scale text-to-image diffusion models. It differs from previous 3D generative models by effectively utilizing web-scale 2D priors while maintaining 3D consistency in a unified model. To bootstrap the training, a lightweight reconstruction model is proposed to instantly produce multi-view Gaussian splat grids for scalable dataset curation. In conjunction with the regular diffusion loss on these grids, a 3D rendering loss is introduced to facilitate 3D coherence across arbitrary views. The compatibility with image diffusion models enables seamless adaptions of numerous techniques for image generation to the 3D realm. Extensive experiments reveal the superiority of DiffSplat in text- and image-conditioned generation tasks and downstream applications. Thorough ablation studies validate the efficacy of each critical design choice and provide insights into the underlying mechanism.

3D生成高斯点云扩散模型多视角一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。