用几何与视觉先验生成单图3D物体,效果更真实一致。
Geometry and Perception Guided Gaussians for Multiview-consistent 3D Generation from a Single Image
- 结合几何与视觉先验初始化高斯分布并优化参数
- 新方法在多视角合成和3D重建上超越现有模型
- 无需额外训练,适合需要高质量3D生成的场景
从单张图像生成逼真的3D物体需兼顾自然外观、3D一致性及对未见区域的多种合理推测。现有方法多依赖微调预训练2D扩散模型或通过快速网络推理直接生成3D信息,但普遍存在多视角不一致和几何细节不足的问题。为此,我们提出一种新方法,无需额外训练即可融合几何与感知先验,重建高细节3D物体。具体地,利用几何先验捕捉粗略3D形状,借助2D预训练扩散模型提供多视角感知信息。随后引入稳定的得分蒸馏采样实现细粒度先验蒸馏,确保知识有效迁移。进一步采用基于重投影的策略强化深度一致性。实验表明,该方法在新视角合成与3D重建任务中优于现有方法,展现出鲁棒且一致的3D生成能力。
原文摘要 · Abstract (English)
Generating realistic 3D objects from single-view images requires natural appearance, 3D consistency, and the ability to capture multiple plausible interpretations of unseen regions. Existing approaches often rely on fine-tuning pretrained 2D diffusion models or directly generating 3D information through fast network inference or 3D Gaussian Splatting, but their results generally suffer from poor multiview consistency and lack geometric detail. To tackle these issues, we present a novel method that seamlessly integrates geometry and perception information without requiring additional model training to reconstruct detailed 3D objects from a single image. Specifically, we incorporate geometry and perception priors to initialize the Gaussian branches and guide their parameter optimization. The geometry prior captures the rough 3D shapes, while the perception prior utilizes the 2D pretrained diffusion model to enhance multiview information. Subsequently, we introduce a stable Score Distillation Sampling for fine-grained prior distillation to ensure effective knowledge transfer. The model is further enhanced by a reprojection-based strategy that enforces depth consistency. Experimental results show that we outperform existing methods on novel view synthesis and 3D reconstruction, demonstrating robust and consistent 3D object generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。