arXiv:2603.19602cs.RO2026-03被引 1

让不同机器人的视觉导航更通用,通过统一几何表示和深度校准提升适应性。

CeRLP: A Cross-embodiment Robot Local Planning Framework for Visual Navigation

  • 将视觉信息抽象为统一几何形式,适配不同尺寸与相机配置的机器人。
  • 引入离线预标定的深度尺度修正,恢复精确的度量深度图。
  • 设计视觉转激光扫描模块,使策略对异构机器人具有鲁棒性,适合多平台部署。

跨体态机器人视觉导航因机器人与摄像头配置差异而困难,现有方法依赖大量数据收集或耗时微调,且常忽略机器人几何结构。本文提出跨体态机器人局部规划框架CeRLP,将视觉信息抽象为统一几何表达,适用于不同物理尺寸、相机参数与类型。该框架引入基于离线预标定的深度估计尺度修正方法,解决单目深度估计的尺度模糊问题,恢复精确度量深度图像。同时,设计视觉到扫描的抽象模块,将异构视觉输入投影为自适应高度的激光扫描,提升策略对多样化机器人的鲁棒性。仿真实验表明,CeRLP优于对比方法,在障碍物规避方面表现优异;真实世界实验进一步验证其在点对点导航与视觉-语言导航任务中的有效性,展现出跨机器人与摄像头配置的强泛化能力。

原文摘要 · Abstract (English)

Visual navigation for cross-embodiment robots is challenging due to variations in robot and camera configurations, which can lead to the failure of navigation tasks. Previous approaches typically rely on collecting massive datasets across different robots, which is highly data-intensive, or fine-tuning models, which is time-consuming. Furthermore, both methods often lack explicit consideration of robot geometry. In this paper, we propose a Cross-embodiment Robot Local Planning (CeRLP) framework for general visual navigation, which abstracts visual information into a unified geometric formulation and applies to heterogeneous robots with varying physical dimensions, camera parameters, and camera types. CeRLP introduces a depth estimation scale correction method that utilizes offline pre-calibration to resolve the scale ambiguity of monocular depth estimation, thereby recovering precise metric depth images. Furthermore, CeRLP designs a visual-to-scan abstraction module that projects varying visual inputs into height-adaptive laser scans, making the policy robust to heterogeneous robots. Experiments in simulation environments demonstrate that CeRLP outperforms comparative methods, validating its robust obstacle avoidance capabilities as a local planner. Additionally, extensive real-world experiments verify the effectiveness of CeRLP in tasks such as point-to-point navigation and vision-language navigation, demonstrating its generalization across varying robot and camera configurations.

视觉导航机器人跨体态局部规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。