arXiv:2502.13519cs.ROcs.AI2025-02ICRA被引 15

用少量专家干预数据训练机器人,让模型学会从反馈中判断动作好坏。

MILE: Model-based Intervention Learning

  • 构建干预模型,从有无干预中提取状态与动作质量信息
  • 仅需少量专家干预,就能在仿真和真实机器人任务中达到良好性能
  • 适合缺乏完整轨迹数据的机器人学习场景

模仿学习在现实控制任务(如机器人)中表现优异,但存在误差累积问题,且通常需要专家提供完整轨迹。尽管已有交互式方法允许专家在必要时介入,但这些方法仅利用干预时段的数据,忽略了非干预时段隐藏的反馈信号。本文提出一种干预建模方法,证明即使在无干预时刻,也能从专家反馈中获取当前状态质量与动作最优性的关键信息。我们仅需少量专家干预即可训练出高效策略。方法在多种离散与连续仿真环境、真实机器人操作任务以及人类实验中均验证有效。视频与代码见https://liralab.usc.edu/mile。

原文摘要 · Abstract (English)

Imitation learning techniques have been shown to be highly effective in real-world control scenarios, such as robotics. However, these approaches not only suffer from compounding error issues but also require human experts to provide complete trajectories. Although there exist interactive methods where an expert oversees the robot and intervenes if needed, these extensions usually only utilize the data collected during intervention periods and ignore the feedback signal hidden in non-intervention timesteps. In this work, we create a model to formulate how the interventions occur in such cases, and show that it is possible to learn a policy with just a handful of expert interventions. Our key insight is that it is possible to get crucial information about the quality of the current state and the optimality of the chosen action from expert feedback, regardless of the presence or the absence of intervention. We evaluate our method on various discrete and continuous simulation environments, a real-world robotic manipulation task, as well as a human subject study. Videos and the code can be found at https://liralab.usc.edu/mile .

模仿学习干预学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。