arXiv:2503.15368cs.ROcs.LG2025-03被引 4

通过动态修正轨迹,减少机器人操作中专家干预次数。

Online Imitation Learning for Manipulation via Decaying Relative Correction through Teleoperation

  • 基于专家提供的空间偏差向量,动态衰减修正动作。
  • 相比传统方法,专家干预率降低30%。
  • 适合需要快速适应新任务的实时操控场景。

遥控机械臂可收集示范数据,用于通过模仿学习训练控制策略。然而,这些方法通常需要大量训练数据才能获得稳健策略或适应新任务。尽管专家反馈能显著提升策略性能,但持续提供反馈对专家而言认知负担重且耗时。为此,我们提出一种缆控遥操作系统,可对策略模型生成的轨迹进行六自由度的空间修正。具体地,我们设计了一种称为衰减相对修正(Decaying Relative Correction, DRC)的方法,基于专家提供的空间偏移向量,该修正具有临时性,从而减少专家所需干预次数。实验表明,DRC相比标准绝对修正方法,将专家干预率降低了30%。此外,将DRC集成到在线模仿学习框架中,能迅速提升如树莓采摘和布料擦拭等操作任务的成功率。

原文摘要 · Abstract (English)

Teleoperated robotic manipulators enable the collection of demonstration data, which can be used to train control policies through imitation learning. However, such methods can require significant amounts of training data to develop robust policies or adapt them to new and unseen tasks. While expert feedback can significantly enhance policy performance, providing continuous feedback can be cognitively demanding and time-consuming for experts. To address this challenge, we propose to use a cable-driven teleoperation system which can provide spatial corrections with 6 degree of freedom to the trajectories generated by a policy model. Specifically, we propose a correction method termed Decaying Relative Correction (DRC) which is based upon the spatial offset vector provided by the expert and exists temporarily, and which reduces the intervention steps required by an expert. Our results demonstrate that DRC reduces the required expert intervention rate by 30\% compared to a standard absolute corrective method. Furthermore, we show that integrating DRC within an online imitation learning framework rapidly increases the success rate of manipulation tasks such as raspberry harvesting and cloth wiping.

机器人操控模仿学习遥操作在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。