从频域出发融合2D/3D先验,提升单图生成3D物体质量
Single Image to Textured 3D Object Generation in Frequency Domain: From Theory to Pipeline

- 在频域中融合2D与3D扩散先验,避免空间域直接拼接的误差
- 在公开和自建数据集上,3D重建质量显著优于现有方法
- 适合需要高细节、多视角一致性的3D内容生成任务
单视图3D重建因信息极度不足而长期具有挑战性。近期基于大规模数据预训练的扩散模型作为2D先验被用于解决该病态问题,但存在颜色偏差和视角不一致的问题。通过使用带有3D标注数据微调的扩散模型作为3D先验可缓解此问题,但其缺乏高频细节,无法通过空间域直接补充2D先验来修复,否则会引入错误的低频2D引导。本文从频域角度重新审视不同扩散先验特性,理论上提出一种多扩散先验在频域中的统一混合优化框架。基于该框架,我们进一步提出Morpheus3D——一种从任意自然未对齐单图生成带纹理3D物体的流程。Morpheus3D通过高通图像提示2D先验引导增强3D先验,有效重建高质量3D物体,同时抑制视角不一致、低频颜色偏差和高频缺失问题。在公共数据集及自建复杂纹理数据集上的定量与定性实验均表明,本方法在生成质量上具有显著提升。
原文摘要 · Abstract (English)
Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extreme lack of information. Recently, diffusion models pre-trained on large-scale datasets served as 2D priors are used to solve the ill-posed task but suffer from color deviation and view inconsistency, which can be curbed by using diffusion models fine-tuned with 3D annotated data served as 3D priors. However, 3D priors lack high-frequency details, which cannot be solved by direct complementation with 2D priors in spatial domain for introducing erroneous low-frequency 2D prior guidance. In this paper, we revisit the characteristics of different diffusion priors from the frequency perspective. Based on our observations, we theoretically present a unified framework of hybrid optimization using multiple diffusion priors in frequency domain. Under this framework, we further propose Morpheus3D, a pipeline of 3D object generation from any single unposed image in the wild. Morpheus3D enhances 3D prior with high-pass image-prompt 2D prior guidance to reconstruct high-quality 3D objects while effectively suppressing view inconsistency, low-frequency color deviation, and high-frequency lacking problems. Both quantitative and qualitative experiments on the public and our collected datasets with complex textures show that our method exhibits significant improvements in generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。