首个跨多年、多海域的水下长期视觉定位数据集,解决定位真值难题。
Long-Term Visual Localization in Dynamic Benthic Environments: A Dataset, Footprint-Based Ground Truth, and Visual Place Recognition Benchmark
- 基于图像足迹构建3D海底覆盖范围,实现精准视觉定位真值标注
- 涵盖5个站点6年回访数据,包含亚分米级相机位姿与立体影像
- 发现现有视觉定位方法在真实水下环境中的召回率显著低于已有基准
长期视觉定位有望降低自主水下航行器(AUV)在光学海底监测中的成本并提升地图质量。然而,由于缺乏可用于基准测试的精选数据集,该领域仍研究不足。此外,地理定位精度有限及图像足迹不明确,要求精确的几何信息以实现准确的真值标注。本文提出一个针对长期水下视觉定位的精选数据集及一种新型真值标注方法。数据集包含来自五个海底参考站点的地理标记AUV影像,最长回访周期达六年,包含原始与色彩校正的立体影像、相机标定参数以及亚分米级注册的相机姿态。据我们所知,这是首个覆盖多个站点和光合带生境的长期水下视觉定位数据集。我们的真值方法通过估计3D海底图像足迹,并将具有重叠足迹的相机视角关联,确保真值链接反映共享视觉内容。基于此数据集与真值,我们对八种前沿视觉位置识别(VPR)方法进行了基准测试,发现其在本数据集上的Recall@K显著低于现有陆地与水下基准。最后,我们将基于足迹的真值与传统的距离阈值真值进行比较,表明在地形起伏大的区域,后者会高估VPR Recall@K。该数据集、真值方法与基准共同为推进动态海底环境中长期视觉定位研究提供基础。
原文摘要 · Abstract (English)
Long-term visual localization has the potential to reduce cost and improve mapping quality in optical benthic monitoring with autonomous underwater vehicles (AUVs). Despite this potential, long-term visual localization in benthic environments remains understudied, primarily due to the lack of curated datasets for benchmarking. Moreover, limited georeferencing accuracy and image footprints necessitate precise geometric information for accurate ground-truthing. In this work, we address these gaps by presenting a curated dataset for long-term visual localization in benthic environments and a novel method to ground-truth visual localization results for near-nadir underwater imagery. Our dataset comprises georeferenced AUV imagery from five benthic reference sites, revisited over periods up to six years, and includes raw and color-corrected stereo imagery, camera calibrations, and sub-decimeter registered camera poses. To our knowledge, this is the first curated underwater dataset for long-term visual localization spanning multiple sites and photic-zone habitats. Our ground-truthing method estimates 3D seafloor image footprints and links camera views with overlapping footprints, ensuring that ground-truth links reflect shared visual content. Building on this dataset and ground truth, we benchmark eight state-of-the-art visual place recognition (VPR) methods and find that Recall@K is significantly lower on our dataset than on established terrestrial and underwater benchmarks. Finally, we compare our footprint-based ground truth to a traditional location-based ground truth and show that distance-threshold ground-truthing can overestimate VPR Recall@K at sites with rugged terrain and altitude variations. Together, the curated dataset, ground-truthing method, and VPR benchmark provide a stepping stone for advancing long-term visual localization in dynamic benthic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。