arXiv:2503.12297cs.RO2025-03被引 2

用人类示范联合学习正反向动力学,让机器人更高效地学会操作技能。

Train Robots in a JIF: Joint Inverse and Forward Dynamics with Human and Robot Demonstrations

  • 联合学习正反向动力学,从多模态人类示范中提取隐状态表示。
  • 仅需少量机器人示范即可高效微调,显著提升数据效率。
  • 支持视觉与触觉融合,适合需要触觉反馈的复杂操作任务。

在大规模机器人示范数据集上进行预训练是学习多样化操作技能的有效方法,但通常受限于收集以机器人为中心的数据的高成本与复杂性,尤其是在需要触觉反馈的任务中。本文提出一种新方法,利用多模态人类示范进行预训练。该方法联合学习逆动力学和前向动力学,以提取隐状态表示,从而学习特定于操作任务的表征。这使得仅需少量机器人示范即可实现高效微调,显著提升数据效率。此外,该方法可结合视觉与触觉等多模态数据,利用隐动力学建模和触觉传感,为基于人类示范的可扩展机器人操作学习开辟新路径。

原文摘要 · Abstract (English)

Pre-training on large datasets of robot demonstrations is a powerful technique for learning diverse manipulation skills but is often limited by the high cost and complexity of collecting robot-centric data, especially for tasks requiring tactile feedback. This work addresses these challenges by introducing a novel method for pre-training with multi-modal human demonstrations. Our approach jointly learns inverse and forward dynamics to extract latent state representations, towards learning manipulation specific representations. This enables efficient fine-tuning with only a small number of robot demonstrations, significantly improving data efficiency. Furthermore, our method allows for the use of multi-modal data, such as combination of vision and touch for manipulation. By leveraging latent dynamics modeling and tactile sensing, this approach paves the way for scalable robot manipulation learning based on human demonstrations.

机器人学习多模态触觉反馈数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。