用模块化设计让机器人精准抓取任意日常物品,实测误差仅2.44厘米。
HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping
- 分模块处理视觉理解与末端执行器控制,结合大模型与仿真训练
- 末端追踪误差降至2.44厘米,比最强基线提升5.5倍
- 适用于办公室、咖啡馆等真实场景,可抓取43至92厘米高物体
在真实环境中对任意物体进行视觉驱动的全身运动操控,需要精确的末端执行器(EE)控制和基于视觉输入(如RGB-D图像)的通用场景理解。现有模仿学习和模拟到现实的方法通过端到端联合学习实现这两方面,难以扩展。本文提出模块化系统HERO,利用大视觉模型实现泛化场景理解,结合仿真训练实现精准的末端执行器控制。核心技术是残差感知的末端跟踪策略:采用逆向运动学将残差目标转为参考轨迹,使用神经前向模型实现精确前向运动学,并引入目标调整与重规划机制。该方法将末端执行器追踪误差降低至2.44厘米,优于最强基线5.5倍。系统可在办公区、咖啡馆等多种真实环境运行,稳定抓取杯子、苹果、玩具等日常物品,操作高度范围为43至92厘米。模块化与端到端对比实验验证了设计有效性。本研究为类人机器人交互日常物体提供了新路径。
原文摘要 · Abstract (English)
Visual loco-manipulation of arbitrary in-the-wild objects requires accurate end-effector (EE) control and a generalizable understanding of the scene from visual inputs (eg, RGB-D images). Existing imitation and sim2real methods jointly learn both these aspects via monolithic end-to-end learning and are thus hard to scale. In this work, we bring to bear the best tools for each of these problems -- large vision models for generalizable scene understanding and simulated training for accurate EE control -- leading to an overall modular loco-manipulation system that exhibits strong generalization. Our core technical innovation is HERO, an accurate residual-aware EE tracking policy made possible by combining classical robotics with machine learning. It uses a) inverse kinematics to convert residual end-effector targets into reference trajectories, b) a learned neural forward model for accurate forward kinematics, and c) goal adjustment and replanning. Together, these innovations reduce the end-effector tracking error to 2.44cm, outperforming the strongest prior method by 5.5x. Our overall system operates in diverse real-world environments, from offices to coffee shops, where the robot reliably grasps various everyday objects (eg, mugs, apples, toys) on surfaces ranging from 43cm to 92cm in height. Systematic modular and end-to-end tests demonstrate the effectiveness of our proposed design. We believe our advances open up new ways of training humanoids to interact with daily objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。