用多视角扩散模型提升3D生成质量,保持各视角一致性。
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
- 引入姿态感知编码器与基于扩散的去噪器,优化低质多视角图像。
- 在多个数据集上实现更高分辨率与更优视角一致性,超越现有方法。
- 适合需要高质量3D建模与多视角一致性的研究者使用。
尽管神经渲染技术取得进展,但受限于高质量3D数据集稀缺及多视角扩散模型的固有缺陷,视图合成与3D模型生成仍局限于低分辨率且多视角一致性不足。本文提出一种新型3D增强流程3DEnhancer,采用多视角潜在扩散模型,在保持多视角一致性的同时增强粗糙3D输入。方法包含姿态感知编码器、基于扩散的去噪器,结合数据增强与多视角注意力模块(含对极线聚合),确保跨视角输出的一致性与高质量。相较于现有视频驱动方法,本模型支持无缝多视角增强,显著提升不同视角间的协同一致性。大量实验表明,3DEnhancer在多视角增强与实例级3D优化任务中均显著优于现有方法。
原文摘要 · Abstract (English)
Despite advances in neural rendering, due to the scarcity of high-quality 3D datasets and the inherent limitations of multi-view diffusion models, view synthesis and 3D model generation are restricted to low resolutions with suboptimal multi-view consistency. In this study, we present a novel 3D enhancement pipeline, dubbed 3DEnhancer, which employs a multi-view latent diffusion model to enhance coarse 3D inputs while preserving multi-view consistency. Our method includes a pose-aware encoder and a diffusion-based denoiser to refine low-quality multi-view images, along with data augmentation and a multi-view attention module with epipolar aggregation to maintain consistent, high-quality 3D outputs across views. Unlike existing video-based approaches, our model supports seamless multi-view enhancement with improved coherence across diverse viewing angles. Extensive evaluations show that 3DEnhancer significantly outperforms existing methods, boosting both multi-view enhancement and per-instance 3D optimization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。