让机器人通过观察学习动作,自动判断哪些动作可行。
Feasibility-aware Imitation Learning from Observation with Multimodal Feedback
- 用机器人动力学模型评估演示动作的可行性
- 实验显示性能提升超3.2倍,动作成功率更高
- 适合人机协作场景中需精准模仿的任务
通过手持设备进行示范来学习机器人控制策略的模仿学习框架受到越来越多关注。然而,由于示范者与机器人在物理特性上的差异,该方法面临两个限制:一是示范数据不包含机器人的动作,二是示范动作可能对机器人不可行。这些限制使策略学习变得困难。为此,我们提出可行性感知的从观察行为克隆(FABCO)。FABCO结合了从观察的行为克隆(利用机器人动力学模型补全动作)与可行性估计。在可行性估计中,使用从机器人执行数据中学习的动力学模型评估示范动作在机器人动力学下的可再现性。估计的可行性用于多模态反馈和可行性感知策略学习,以优化示范动作并学习稳健策略。多模态反馈通过示范者的视觉和触觉感知提供可行性信息,促进更可行的示范动作。可行性感知策略学习降低对机器人不可行动作的影响,实现机器人可稳定执行的策略学习。我们在15名参与者上针对两个任务进行了实验,结果表明,相比无可行性反馈的情况,FABCO将模仿学习性能提升了3.2倍以上。
原文摘要 · Abstract (English)
Imitation learning frameworks that learn robot control policies from demonstrators' motions via hand-mounted demonstration interfaces have attracted increasing attention. However, due to differences in physical characteristics between demonstrators and robots, this approach faces two limitations: i) the demonstration data do not include robot actions, and ii) the demonstrated motions may be infeasible for robots. These limitations make policy learning difficult. To address them, we propose Feasibility-Aware Behavior Cloning from Observation (FABCO). FABCO integrates behavior cloning from observation, which complements robot actions using robot dynamics models, with feasibility estimation. In feasibility estimation, the demonstrated motions are evaluated using a robot-dynamics model, learned from the robot's execution data, to assess reproducibility under the robot's dynamics. The estimated feasibility is used for multimodal feedback and feasibility-aware policy learning to improve the demonstrator's motions and learn robust policies. Multimodal feedback provides feasibility through the demonstrator's visual and haptic senses to promote feasible demonstrated motions. Feasibility-aware policy learning reduces the influence of demonstrated motions that are infeasible for robots, enabling the learning of policies that robots can execute stably. We conducted experiments with 15 participants on two tasks and confirmed that FABCO improves imitation learning performance by more than 3.2 times compared to the case without feasibility feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。