arXiv:2409.13607cs.RO2024-09被引 4

人类用标记物引导机器人学习,减少环境干扰导致的误解。

RECON: Reducing Causal Confusion with Human-Placed Markers

  • 人类在关键物体上贴标记物,机器人通过追踪标记学习任务。
  • 使用标记数据训练状态嵌入,使机器人自动忽略无关信息。
  • 大幅减少演示次数,降低人工教学时间,适合人机协作场景。

模仿学习让机器人从人类示范中学习新任务,但存在因果混淆问题:机器人观察到的任务相关信息与环境干扰(如杂乱、光照变化)混在一起,难以判断哪些特征重要,导致学习失败。本文提出RECON框架,让人类教师在任务相关物体上主动放置小型轻量标记物。示范过程中,标记物位置被实时追踪,利用这些离线标记数据训练出与标记读数相关的任务相关状态嵌入。该嵌入将机器人观测映射到与标记信号强相关的潜在状态,使其能自动过滤无关信息,基于标记数据学习的关键特征做出决策。仿真和真实机器人实验表明,该方法有效缓解因果混淆,显著减少完成任务所需的示范数量,从而降低整体教学时间。

原文摘要 · Abstract (English)

Imitation learning enables robots to learn new tasks from human examples. One fundamental limitation while learning from humans is causal confusion. Causal confusion occurs when the robot's observations include both task-relevant and extraneous information: for instance, a robot's camera might see not only the intended goal, but also clutter and changes in lighting within its environment. Because the robot does not know which aspects of its observations are important a priori, it often misinterprets the human's examples and fails to learn the desired task. To address this issue, we highlight that -- while the robot learner may not know what to focus on -- the human teacher does. In this paper we propose that the human proactively marks key parts of their task with small, lightweight beacons. Under our framework (RECON) the human attaches these beacons to task-relevant objects before providing demonstrations: as the human shows examples of the task, beacons track the position of marked objects. We then harness this offline beacon data to train a task-relevant state embedding. Specifically, we embed the robot's observations to a latent state that is correlated with the measured beacon readings: in practice, this causes the robot to autonomously filter out extraneous observations and make decisions based on features learned from the beacon data. Our simulations and a real robot experiment suggest that this framework for human-placed beacons mitigates causal confusion. Indeed, we find that using RECON significantly reduces the number of demonstrations needed to convey the task, lowering the overall time required for human teaching. See videos here: https://youtu.be/oy85xJvtLSU

机器人学习模仿学习人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。