跨视图补全模型可零样本估计图像对应关系,效果优于传统特征。
Cross-View Completion Models are Zero-shot Correspondence Estimators
- 利用交叉注意力图捕捉跨视图对应关系
- 在零样本匹配与深度估计任务中表现优异
- 适合做无监督视觉对齐或几何理解的研究者
本文从自监督对应学习的角度重新审视跨视图补全学习。分析表明,跨视图补全模型中的交叉注意力图比编码器或解码器特征提取的其他相关性更能有效捕捉对应关系。通过在零样本匹配、基于学习的几何匹配及多帧深度估计任务上的评估,验证了交叉注意力图的有效性。项目页面见 https://cvlab-kaist.github.io/ZeroCo/。
原文摘要 · Abstract (English)
In this work, we explore new perspectives on cross-view completion learning by drawing an analogy to self-supervised correspondence learning. Through our analysis, we demonstrate that the cross-attention map within cross-view completion models captures correspondence more effectively than other correlations derived from encoder or decoder features. We verify the effectiveness of the cross-attention map by evaluating on both zero-shot matching and learning-based geometric matching and multi-frame depth estimation. Project page is available at https://cvlab-kaist.github.io/ZeroCo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。