提出新数据集C3,解决照片与平面图跨视角跨模态匹配难题。
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- 构建3D场景并手动对齐平面图,生成90万张图像-平面图对应关系。
- 在新数据上训练使模型误差降低34%,相机位姿估计准确率显著提升。
- 适合研究跨模态几何理解、视觉定位与建筑信息建模的学者参考。
如DUSt3R等几何模型虽在图片对几何理解方面取得进展,但在输入视角(如航拍与地面)或模态(如照片与抽象绘图)差异较大时表现不佳。本文聚焦于地面照片与平面图之间的对应关系预测问题。现有数据集要么缺乏多模态(VIGOR),要么缺少对应标注(WAFFLE)。为此,我们构建了新数据集C3:先通过结构光恢复互联网照片集合中的3D场景,再将重建结果人工对齐来自网络的平面图,从而获得图像与平面图间的对应关系。C3包含597个场景中90,000对平面图与照片,共1.53亿像素级对应关系及85,000个相机位姿。我们发现当前最优对应模型在此任务上表现仍弱。在本数据集上训练后,最佳方法在均方根误差(RMSE)上提升34%。我们还利用预测对应关系估计相机位姿,并以召回率评估性能。最后,指出了跨模态几何推理中的开放挑战,该数据集旨在推动相关研究发展。
原文摘要 · Abstract (English)
Geometric models like DUSt3R have shown great advances in understanding the geometry of a scene from pairs of photos. However, they fail when the inputs are from vastly different viewpoints (e.g., aerial vs. ground) or modalities (e.g., photos vs. abstract drawings) compared to what was observed during training. This paper addresses a challenging version of this problem: predicting correspondences between ground-level photos and floor plans. Current datasets for joint photo-floor plan reasoning are limited, either lacking in varying modalities (VIGOR) or lacking in correspondences (WAFFLE). To address these limitations, we introduce a new dataset, C3, created by first reconstructing a number of scenes in 3D from Internet photo collections via structure-from-motion, then manually registering the reconstructions to floor plans gathered from the Internet, from which we can derive correspondences between images and floor plans. C3 contains 90K paired floor plans and photos across 597 scenes with 153M pixel-level correspondences and 85K camera poses. We find that state-of-the-art correspondence models struggle on this task. By training on our new data, we can improve on the best performing method by 34% in RMSE. We also use the predicted correspondences to estimate camera poses and evaluate performance using recall metrics. Lastly, we identify open challenges in cross-modal geometric reasoning that our dataset aims to help address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。