系统梳理跨视角定位与生成任务的进展与挑战。
Cross-view Localization and Synthesis -- Datasets, Challenges and Opportunities
- 将地表图像与航拍图匹配定位,或从航拍图生成地表视图。
- 使用CNN/ViT提取特征进行跨视角检索,用GAN/扩散模型生成图像。
- 适合研究视觉理解、自动驾驶与城市规划的学者参考。
跨视角定位与合成是跨视角视觉理解中的两个基础任务,涉及航拍(卫星或航空)与地面影像数据集。由于在自动驾驶、城市规划和增强现实等领域的广泛应用,该领域日益受到关注。跨视角定位旨在根据航拍图像信息估计地面图像的地理坐标;而跨视角合成则尝试基于航拍信息生成地面视图。两者均因视角、分辨率和遮挡差异显著而面临挑战,这些差异广泛存在于跨视角数据集中。近年来,得益于大规模数据集和新方法的推动,该领域取得快速进展。通常,跨视角定位被建模为图像检索问题,通过卷积神经网络(CNN)或视觉变换器(ViTs)提取特征,在瓦片化航拍图像中匹配地面特征。跨视角合成则多采用生成对抗网络(GANs)或扩散模型,根据航拍信息生成地面视图。本文全面综述了跨视角定位与合成的最新进展,回顾常用数据集,强调关键挑战,系统梳理前沿技术。同时分析现有局限,提供对比分析,并展望未来研究方向。项目主页见:https://github.com/GDAOSU/Awesome-Cross-View-Methods。
原文摘要 · Abstract (English)
Cross-view localization and synthesis are two fundamental tasks in cross-view visual understanding, which deals with cross-view datasets: overhead (satellite or aerial) and ground-level imagery. These tasks have gained increasing attention due to their broad applications in autonomous navigation, urban planning, and augmented reality. Cross-view localization aims to estimate the geographic position of ground-level images based on information provided by overhead imagery while cross-view synthesis seeks to generate ground-level images based on information from the overhead imagery. Both tasks remain challenging due to significant differences in viewing perspective, resolution, and occlusion, which are widely embedded in cross-view datasets. Recent years have witnessed rapid progress driven by the availability of large-scale datasets and novel approaches. Typically, cross-view localization is formulated as an image retrieval problem where ground-level features are matched with tiled overhead images feature, extracted by convolutional neural networks (CNNs) or vision transformers (ViTs) for cross-view feature embedding. Cross-view synthesis, on the other hand, seeks to generate ground-level views based on information from overhead imagery, generally using generative adversarial networks (GANs) or diffusion models. This paper presents a comprehensive survey of advances in cross-view localization and synthesis, reviewing widely used datasets, highlighting key challenges, and providing an organized overview of state-of-the-art techniques. Furthermore, it discusses current limitations, offers comparative analyses, and outlines promising directions for future research. We also include the project page via https://github.com/GDAOSU/Awesome-Cross-View-Methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。