arXiv:2512.11173cs.RO2025-12被引 1

仅用摄像头图像实现物体类别级精准定位,无需深度或地图信息

Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance

  • 基于视觉模仿学习,结合目标图像与文本提示进行定位决策
  • 在未见物体上达成74.6%边缘对齐与89.4%物体对齐成功率
  • 适用于真实复杂光照与背景,适合移动操作机器人部署

实现移动操作机器人基座的精确定位对后续操作至关重要。现有大多数基于RGB的导航系统仅能提供米级粗略精度,难以满足最后米级操作定位需求,导致操作策略无法在训练分布内执行,频繁失败。本文提出一种面向类别级最后米级导航的物体中心模仿学习框架,使四足移动操作机器人仅依赖机载摄像头的RGB观测,即可实现操作就绪定位。方法以目标图像、多视角RGB观测和指定目标物体的文本提示为输入,通过语言驱动分割模块与空间评分矩阵解码器实现显式物体定位与相对姿态推理。仅使用类别中单个实例的真实世界数据,系统可泛化至不同环境中的未见物体实例,即使在挑战性光照与背景条件下仍有效。为全面评估,引入两项指标:边缘对齐(基于真实朝向)与物体对齐(评估机器人视觉朝向目标程度)。实验结果表明,在未见目标物体上,该策略在边缘对齐任务中达到74.58%成功率,在物体对齐任务中达到89.42%成功率。结果证明,无需深度、激光雷达或地图先验,即可实现类别级高精度最后米级导航,为统一移动操作提供了可扩展路径。

原文摘要 · Abstract (English)

Achieving precise positioning of the mobile manipulator's base is essential for successful manipulation actions that follow. Most of the RGB-based navigation systems only guarantee coarse, meter-level accuracy, making them less suitable for the precise positioning phase of mobile manipulation. This gap prevents manipulation policies from operating within the distribution of their training demonstrations, resulting in frequent execution failures. We address this gap by introducing an object-centric imitation learning framework for last-meter navigation, enabling a quadruped mobile manipulator robot to achieve manipulation-ready positioning using only RGB observations from its onboard cameras. Our method conditions the navigation policy on three inputs: goal images, multi-view RGB observations from the onboard cameras, and a text prompt specifying the target object. A language-driven segmentation module and a spatial score-matrix decoder then supply explicit object grounding and relative pose reasoning. Using real-world data from a single object instance within a category, the system generalizes to unseen object instances across diverse environments with challenging lighting and background conditions. To comprehensively evaluate this, we introduce two metrics: an edge-alignment metric, which uses ground truth orientation, and an object-alignment metric, which evaluates how well the robot visually faces the target. Under these metrics, our policy achieves 74.58% success in edge-alignment and 89.42% success in object-alignment when positioning relative to unseen target objects. These results show that precise last-meter navigation can be achieved at a category-level without depth, LiDAR, or map priors, enabling a scalable pathway toward unified mobile manipulation. Project page: https://rpm-lab-umn.github.io/category-level-last-meter-nav/

导航模仿学习移动操作视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。