让机器人看懂动作坐标系,提升视觉操作的鲁棒性
AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

- 用相机参数和末端姿态渲染坐标轴,生成动作方向提示图
- 在真实与仿真环境中均显著提升抓取成功率,泛化能力更强
- 适合做通用视觉-动作策略的开发者,尤其关注迁移性能
通过大规模行为克隆训练的视觉-动作操作策略虽具备较强的语义场景理解能力,但在分布外情况下常无法可靠执行底层动作。例如,在相同场景布局、相机视角和光照下,当物体位于未见过的位置时,性能仍会显著下降。我们认为这一差距源于动作理解不足——即无法在图像空间中解释机器人基座坐标系的动作含义。为此,我们提出AxisGuide,一种轻量级引导方法,连接语义场景理解与动作坐标解释。利用相机参数和末端执行器位姿,AxisGuide在每个相机视角中渲染机器人基座坐标轴,并通过少量提示通道在RGB观测中显式可视化+x、+y、+z运动在图像中的意义。在LIBERO仿真与真实世界环境中的大量评估表明,AxisGuide带来显著性能提升与更好泛化能力,验证了显式动作坐标提示在学习可靠、可迁移的通用视觉-动作策略中的有效性。
原文摘要 · Abstract (English)
Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliably execute correct low-level actions under distribution shifts. For example, even in a simple pickup task with identical scene layouts, camera viewpoints, and illumination, performance can degrade substantially when the object is placed at unseen locations. We argue that this gap arises from insufficient action understanding, namely the inability to interpret the robot's base-frame action coordinate system in image space. To address this issue, we introduce AxisGuide, a lightweight guidance method that bridges semantic scene understanding and action-coordinate interpretation. Using camera parameters and end-effector poses, AxisGuide renders the robot base-frame axes in each camera view and augments RGB observations with a small set of cue channels that explicitly visualize the meaning of the +x, +y, and +z motions in image space. Extensive evaluations in both the LIBERO simulation and real-world environments demonstrate that AxisGuide yields substantial performance gains and improved generalization, highlighting the effectiveness of explicit action-coordinate cues for learning reliable and transferable generalist visuomotor policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。