arXiv:2510.04171cs.RO2025-10

从俯视图直接学习机器人基座最优位置,提升抓取效率。

VBM-NET: Visual Base Pose Learning for Mobile Manipulation using Equivariant TransporterNet and GNNs

  • 用等变TransporterNet捕捉空间对称性,高效生成候选基座位姿
  • 通过GNN和强化学习在候选位姿中选出最优解,计算耗时更少
  • 可在仿真训练后直接部署到真实机器人,具备强泛化能力

在移动操作中,选择最优的移动基座位姿对成功抓取物体至关重要。以往方法依赖精确的状态信息(如物体位姿、环境模型),通过经典规划或基于状态的策略求解。本文研究直接从场景的俯视正射投影图像中进行基座位姿规划,该图像提供全局场景概览并保留空间结构。我们提出VBM-NET,一种基于学习的方法,利用此类俯视图进行基座位姿选择。采用等变TransporterNet挖掘空间对称性,高效学习抓取候选基座位姿;进一步使用图神经网络表示可变数量的候选位姿,并通过强化学习从中选出最优位姿。实验表明,VBM-NET在显著更少的计算时间内即可获得与经典方法相当的结果。此外,通过仿真到现实的迁移验证,成功将仿真中训练的策略部署到真实移动操作任务中。

原文摘要 · Abstract (English)

In Mobile Manipulation, selecting an optimal mobile base pose is essential for successful object grasping. Previous works have addressed this problem either through classical planning methods or by learning state-based policies. They assume access to reliable state information, such as the precise object poses and environment models. In this work, we study base pose planning directly from top-down orthographic projections of the scene, which provide a global overview of the scene while preserving spatial structure. We propose VBM-NET, a learning-based method for base pose selection using such top-down orthographic projections. We use equivariant TransporterNet to exploit spatial symmetries and efficiently learn candidate base poses for grasping. Further, we use graph neural networks to represent a varying number of candidate base poses and use Reinforcement Learning to determine the optimal base pose among them. We show that VBM-NET can produce comparable solutions to the classical methods in significantly less computation time. Furthermore, we validate sim-to-real transfer by successfully deploying a policy trained in simulation to real-world mobile manipulation.

移动操作位姿规划等变网络仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。