用几何信息提升多视角图像生成的一致性与细节
GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
- 通过深度图、法线图等提取共享几何结构
- 生成图像在多视角间保持形状一致且细节丰富
- 适合3D重建、虚拟现实等需要高保真多视图的场景
多视角图像生成在计算机视觉中具有重要应用价值,尤其在3D重建、虚拟现实和增强现实领域。现有方法多基于单图扩展,面临跨视角一致性难维持和高分辨率生成计算成本高的问题。为此,我们提出几何引导的多视角扩散模型(GeoMVD),通过提取多视角几何信息并调节几何特征强度,实现视角间一致且细节丰富的图像生成。具体地,设计了多视角几何信息提取模块,利用深度图、法线图和前景分割掩码构建共享几何结构,确保不同视角间形状与结构一致。为增强生成过程中的一致性与细节恢复,提出解耦式几何增强注意力机制,强化关键几何细节的特征关注,提升整体图像质量与细节保留能力。此外,采用自适应学习策略优化空间关系与视图间视觉连贯性,确保生成结果真实可信。模型还引入迭代细化过程,通过多阶段生成逐步提升输出质量。最后,提出动态几何信息强度调节机制,自适应控制几何数据影响,平衡生成质量与自然性。更多细节见项目页:https://sobeymil.github.io/GeoMVD.com。
原文摘要 · Abstract (English)
Multi-view image generation holds significant application value in computer vision, particularly in domains like 3D reconstruction, virtual reality, and augmented reality. Most existing methods, which rely on extending single images, face notable computational challenges in maintaining cross-view consistency and generating high-resolution outputs. To address these issues, we propose the Geometry-guided Multi-View Diffusion Model, which incorporates mechanisms for extracting multi-view geometric information and adjusting the intensity of geometric features to generate images that are both consistent across views and rich in detail. Specifically, we design a multi-view geometry information extraction module that leverages depth maps, normal maps, and foreground segmentation masks to construct a shared geometric structure, ensuring shape and structural consistency across different views. To enhance consistency and detail restoration during generation, we develop a decoupled geometry-enhanced attention mechanism that strengthens feature focus on key geometric details, thereby improving overall image quality and detail preservation. Furthermore, we apply an adaptive learning strategy that fine-tunes the model to better capture spatial relationships and visual coherence between the generated views, ensuring realistic results. Our model also incorporates an iterative refinement process that progressively improves the output quality through multiple stages of image generation. Finally, a dynamic geometry information intensity adjustment mechanism is proposed to adaptively regulate the influence of geometric data, optimizing overall quality while ensuring the naturalness of generated images. More details can be found on the project page: https://sobeymil.github.io/GeoMVD.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。