arXiv:2412.04470cs.CV2024-12CVPR被引 15

一秒钟生成高质量3D模型,速度远超此前方法。

Turbo3D: Ultra-fast Text-to-3D Generation

  • 四步四视角扩散生成+潜空间重建,实现极速推理。
  • 生成3D资产仅需0.9秒,且质量优于现有基线。
  • 适合需要快速原型设计或实时交互的开发者。

我们提出Turbo3D,一种超高速文本到3D生成系统,可在不到1秒内生成高质量的高斯点云资产。Turbo3D采用快速的四步四视角扩散生成器与高效的前馈高斯重构器,均在潜空间中运行。该四步四视角生成器是通过一种新型双教师方法蒸馏得到的学生模型,促使学生从多视角教师学习视图一致性,从单视角教师学习照片级真实感。通过将高斯重构器的输入从像素空间转移到潜空间,我们消除了额外的图像解码时间,并将Transformer序列长度减半以实现最大效率。与先前基线相比,本方法在生成质量上表现更优,同时运行时间仅为它们的一小部分。

原文摘要 · Abstract (English)

We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multi-view teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor's inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime.

3D生成扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。