让2D扩散模型生成更立体的3D物体,通过3D反馈提升几何一致性。
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
- 在去噪过程中动态构建3D结构并反馈到2D模型中
- 在Instant3D等模型上显著提升3D几何质量,生成更真实物体
- 无需训练即可适配多类任务,适合3D生成初学者和研究者
多视角图像扩散模型已显著推动开放域3D物体生成。然而,现有方法多依赖缺乏3D先验的2D网络架构,导致几何一致性不足。为此,我们提出3D-Adapter,一个可插拔模块,用于向预训练图像扩散模型注入3D几何感知能力。核心思想是3D反馈增强:在采样循环的每一步,3D-Adapter将多视角中间特征解码为一致的3D表示,再渲染出RGBD视图并重新编码,通过特征相加方式增强基线模型。我们研究了两种变体:基于高斯点云的快速前馈版本,以及利用神经场和网格的免训练通用版本。大量实验表明,3D-Adapter不仅显著提升Instant3D和Zero123++等文本到多视角模型的几何质量,还能使普通文生图Stable Diffusion实现高质量3D生成。此外,我们在文生3D、图生3D、文生纹理和文生虚拟角色任务中均取得高质量结果,展现其广泛应用潜力。
原文摘要 · Abstract (English)
Multi-view image diffusion models have significantly advanced open-domain 3D object generation. However, most existing models rely on 2D network architectures that lack inherent 3D biases, resulting in compromised geometric consistency. To address this challenge, we introduce 3D-Adapter, a plug-in module designed to infuse 3D geometry awareness into pretrained image diffusion models. Central to our approach is the idea of 3D feedback augmentation: for each denoising step in the sampling loop, 3D-Adapter decodes intermediate multi-view features into a coherent 3D representation, then re-encodes the rendered RGBD views to augment the pretrained base model through feature addition. We study two variants of 3D-Adapter: a fast feed-forward version based on Gaussian splatting and a versatile training-free version utilizing neural fields and meshes. Our extensive experiments demonstrate that 3D-Adapter not only greatly enhances the geometry quality of text-to-multi-view models such as Instant3D and Zero123++, but also enables high-quality 3D generation using the plain text-to-image Stable Diffusion. Furthermore, we showcase the broad application potential of 3D-Adapter by presenting high quality results in text-to-3D, image-to-3D, text-to-texture, and text-to-avatar tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。