arXiv:2603.01475cs.CV2026-03被引 3

构建首个大规模自然环境跨模态基准,支持定位与深度估计

WildCross: A Cross-Modal Large Scale Benchmark for Place Recognition and Metric Depth Estimation in Natural Environments

  • 采集47.6万帧带稠密深度的序列图像,配准激光雷达与位姿
  • 在自然场景中验证多模态定位与度量深度估计性能
  • 适合研究野外机器人感知与跨模态学习的学者使用

近年来,对非结构化自然环境中机器人解决方案的需求显著增长,同时人们对融合2D与3D场景理解的兴趣也日益浓厚。然而,现有机器人数据集主要来自结构化的城市环境,难以应对复杂、非结构化自然场景的挑战。为此,我们提出WildCross,一个面向大规模自然环境中场景识别与度量深度估计的跨模态基准。WildCross包含超过47.6万帧连续的RGB图像,配有半稠密深度与表面法向标注,每帧均对齐精确的6DoF位姿及同步的密集激光雷达子地图。我们在视觉、激光雷达和跨模态场景识别以及度量深度估计方面进行了全面实验,证明了WildCross作为多模态机器人感知任务的挑战性基准的价值。代码与数据集可通过https://csiro-robotics.github.io/WildCross获取。

原文摘要 · Abstract (English)

Recent years have seen a significant increase in demand for robotic solutions in unstructured natural environments, alongside growing interest in bridging 2D and 3D scene understanding. However, existing robotics datasets are predominantly captured in structured urban environments, making them inadequate for addressing the challenges posed by complex, unstructured natural settings. To address this gap, we propose WildCross, a cross-modal benchmark for place recognition and metric depth estimation in large-scale natural environments. WildCross comprises over 476K sequential RGB frames with semi-dense depth and surface normal annotations, each aligned with accurate 6DoF poses and synchronized dense lidar submaps. We conduct comprehensive experiments on visual, lidar, and cross-modal place recognition, as well as metric depth estimation, demonstrating the value of WildCross as a challenging benchmark for multi-modal robotic perception tasks. We provide access to the code repository and dataset at https://csiro-robotics.github.io/WildCross.

场景识别深度估计跨模态自然环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。