新基准揭示视觉模型在自然环境中的感知短板
Cross-Modal Benchmarking for Robotic Perception in Natural Environments

- 构建跨模态基准WildCross,含476K帧序列与精准位姿标注
- 现有视觉模型在自然场景深度估计上误差高达18.7%
- 适合野外机器人、多模态感知研究者参考
自然环境对机器人感知系统构成复杂挑战。当前模型,尤其是视觉基础模型,主要在结构化城市环境中训练,导致在野外任务中感知能力不足。本文通过我们近期发布的WildCross基准,展示了现有模型的局限性。该基准是大规模自然环境中用于场景识别和度量深度估计的跨模态基准,包含超过476,000帧的连续RGB图像,配有半密集深度与表面法向标注,每帧均与精确的6DoF位姿及同步的稠密激光雷达子图对齐。本文对基准结果进行了扩展分析,重点开展度量深度估计实验。代码与数据集可通过https://csiro-robotics.github.io/WildCross获取。
原文摘要 · Abstract (English)
Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trained on structured, urban environments leading to weaknesses in their perception for field robotics tasks. We showcase the limitations of current models using our recently released WildCross benchmark, a new cross-modal benchmark for place recognition and metric depth estimation in large-scale natural environments. WildCross comprises over 476K sequential RGB frames with semi-dense depth and surface normal annotations, each aligned with accurate 6DoF pose and synchronized dense lidar submaps. In this work, we provide an expanded analysis of the benchmark results from the recent WildCross benchmark, with particular emphasis on expanded metric depth estimation experiments. Access to the code repository and dataset for this work can be found at https://csiro-robotics.github.io/WildCross.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。