用扩散模型生成更准更真的立体图像,还能无监督匹配
Towards Open-World Generation of Stereo Images and Unsupervised Matching
- 用视差感知坐标嵌入和扭曲图像引导扩散过程
- 在11个数据集上实现图像质量和视差一致性双优
- 适合做XR、自动驾驶立体视觉的开发者
立体图像在扩展现实(XR)设备、自动驾驶和机器人等领域至关重要。然而,由于双摄系统校准要求高,且获取精确稠密视差图复杂,高质量立体图像的获取仍具挑战。现有生成方法通常只关注视觉质量或几何准确性,难以兼顾两者。本文提出GenStereo,一种基于扩散模型的方法,通过两个关键创新实现突破:(1) 在扩散过程中引入视差感知坐标嵌入与扭曲输入图像,提升立体对齐精度;(2) 设计自适应融合机制,智能结合生成图像与扭曲图像,增强真实感与视差一致性。在11个多样化的立体图像数据集上进行充分训练后,GenStereo展现出强大的泛化能力,在立体图像生成与无监督立体匹配任务中均达到当前最优性能。
原文摘要 · Abstract (English)
Stereo images are fundamental to numerous applications, including extended reality (XR) devices, autonomous driving, and robotics. Unfortunately, acquiring high-quality stereo images remains challenging due to the precise calibration requirements of dual-camera setups and the complexity of obtaining accurate, dense disparity maps. Existing stereo image generation methods typically focus on either visual quality for viewing or geometric accuracy for matching, but not both. We introduce GenStereo, a diffusion-based approach, to bridge this gap. The method includes two primary innovations (1) conditioning the diffusion process on a disparity-aware coordinate embedding and a warped input image, allowing for more precise stereo alignment than previous methods, and (2) an adaptive fusion mechanism that intelligently combines the diffusion-generated image with a warped image, improving both realism and disparity consistency. Through extensive training on 11 diverse stereo datasets, GenStereo demonstrates strong generalization ability. GenStereo achieves state-of-the-art performance in both stereo image generation and unsupervised stereo matching tasks. Project page is available at https://qjizhi.github.io/genstereo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。