修正视角生成缺陷,提升多图条件下的视图一致性
Fixing the Perspective: A Critical Examination of Zero-1-to-3
- 重设计交叉注意力机制,解决多视角条件信息处理偏差
- 改进架构支持同时利用多个输入图像生成新视图
- 为高质量3D图像生成提供更可靠的视觉一致性基础
新颖视图合成是图像到3D生成的核心挑战,需从一组条件图像及其相对位姿生成目标视图。尽管近期方法如Zero-1-to-3利用条件潜空间扩散模型取得了良好效果,但在处理多张条件图像时仍存在视图生成不一致、不准确的问题。本文深入分析Zero-1-to-3在扩散2D条件UNet的空间变换器中交叉注意力机制的实现,发现其理论框架与实际实现间存在关键偏差,尤其体现在图像条件上下文的处理上。为此提出两项改进:(1)修正后的交叉注意力实现,使条件信息得以有效利用;(2)增强型架构,可同时融合多个条件视图。理论分析与初步结果表明,该方法有望显著提升新视图合成的一致性与准确性。
原文摘要 · Abstract (English)
Novel view synthesis is a fundamental challenge in image-to-3D generation, requiring the generation of target view images from a set of conditioning images and their relative poses. While recent approaches like Zero-1-to-3 have demonstrated promising results using conditional latent diffusion models, they face significant challenges in generating consistent and accurate novel views, particularly when handling multiple conditioning images. In this work, we conduct a thorough investigation of Zero-1-to-3's cross-attention mechanism within the Spatial Transformer of the diffusion 2D-conditional UNet. Our analysis reveals a critical discrepancy between Zero-1-to-3's theoretical framework and its implementation, specifically in the processing of image-conditional context. We propose two significant improvements: (1) a corrected implementation that enables effective utilization of the cross-attention mechanism, and (2) an enhanced architecture that can leverage multiple conditional views simultaneously. Our theoretical analysis and preliminary results suggest potential improvements in novel view synthesis consistency and accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。