用视线和运动瓶颈提升机器人抓取技能的复用性。
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
- 结合视线与运动瓶颈信息,增强技能泛化能力。
- 在物体位置变化时仍保持高成功率,优于现有方法。
- 完全数据驱动训练,适合需要灵活操作的机器人场景。
能够执行多样化物体操作的自主智能体应具备广泛且可复用的操作技能。尽管深度学习使机器人复制人类遥控操作的灵巧性成为可能,但将习得技能泛化到未见场景仍是重大挑战。本文提出一种新算法GazeBot,通过利用视线信息和运动瓶颈这两个关键特征,在不牺牲灵巧性或反应速度的前提下,显著提升已学动作的复用性。实验表明,当物体位置及末端执行器位姿与示范数据存在差异时,GazeBot相较于前沿模仿学习方法仍能保持更高成功率。该方法训练过程完全数据驱动,只需提供带有视线数据的示范数据集即可。视频与代码已公开于https://crumbyrobotics.github.io/gazebot。
原文摘要 · Abstract (English)
Autonomous agents capable of diverse object manipulations should be able to acquire a wide range of manipulation skills with high reusability. Although advances in deep learning have made it increasingly feasible to replicate the dexterity of human teleoperation in robots, generalizing these acquired skills to previously unseen scenarios remains a significant challenge. In this study, we propose a novel algorithm, Gaze-based Bottleneck-aware Robot Manipulation (GazeBot), which enables high reusability of learned motions without sacrificing dexterity or reactivity. By leveraging gaze information and motion bottlenecks, both crucial features for object manipulation, GazeBot achieves high success rates compared with state-of-the-art imitation learning methods, particularly when the object positions and end-effector poses differ from those in the provided demonstrations. Furthermore, the training process of GazeBot is entirely data-driven once a demonstration dataset with gaze data is provided. Videos and code are available at https://crumbyrobotics.github.io/gazebot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。