让机器人学习人类示范时,先判断动作能否执行,提升成功率。
Feasibility-aware Imitation Learning from Observations through a Hand-mounted Demonstration Interface
- 用机器人的动力学模型评估人类示范的可行性
- 可行性反馈使演示更准确,加权训练提升策略鲁棒性
- 适合希望高效学习真实操作的机器人开发者
通过示范接口进行模仿学习,旨在从直观的人类示范中学习机器人自动化策略。然而,由于人与机器人运动特性差异,人类专家可能无意中示范出机器人无法执行的动作。本文提出可行性感知的行为克隆方法(FABCO):利用机器人预训练的前向和逆向动力学模型评估每个示范的可行性,并将该信息以视觉反馈形式提供给示范者,引导其优化示范动作。在策略学习阶段,估计的可行性作为示范数据的权重,提升了数据效率和策略鲁棒性。我们在移液管插入任务中验证了FABCO的有效性,四名参与者评估了可行性反馈与加权学习的影响。同时使用NASA任务负荷指数(NASA-TLX)评估了带视觉反馈示范带来的工作负荷。
原文摘要 · Abstract (English)
Imitation learning through a demonstration interface is expected to learn policies for robot automation from intuitive human demonstrations. However, due to the differences in human and robot movement characteristics, a human expert might unintentionally demonstrate an action that the robot cannot execute. We propose feasibility-aware behavior cloning from observation (FABCO). In the FABCO framework, the feasibility of each demonstration is assessed using the robot's pre-trained forward and inverse dynamics models. This feasibility information is provided as visual feedback to the demonstrators, encouraging them to refine their demonstrations. During policy learning, estimated feasibility serves as a weight for the demonstration data, improving both the data efficiency and the robustness of the learned policy. We experimentally validated FABCO's effectiveness by applying it to a pipette insertion task involving a pipette and a vial. Four participants assessed the impact of the feasibility feedback and the weighted policy learning in FABCO. Additionally, we used the NASA Task Load Index (NASA-TLX) to evaluate the workload induced by demonstrations with visual feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。