arXiv:2510.19356cs.RO2025-10被引 1

提出多步一致性捷径模型,加速机器人模仿学习推理且保持高精度。

Imitation Learning Policy based on Multi-Step Consistent Integration Shortcut Model

  • 基于捷径模型扩展多步一致性损失,分拆单步损失为多步以提升性能。
  • 在仿真与真实场景中验证,单步推理速度显著提升,性能优于现有方法。
  • 适合追求高速高精度的机器人控制任务,尤其适用于实时系统部署。

流匹配方法的广泛应用极大地推动了机器人模仿学习的发展,但这些方法普遍存在推理时间过长的问题。为此,研究人员提出了蒸馏法和一致性方法,但其性能仍难以超越原始扩散模型和流匹配模型。本文提出一种基于多步一致集成的单步捷径方法,用于机器人模仿学习。为平衡推理速度与性能,我们在捷径模型基础上扩展多步一致性损失,将单步损失分解为多步损失,并提升单步推理性能。其次,为解决多步损失与原始流匹配损失优化不稳定的难题,我们提出自适应梯度分配方法,增强学习过程的稳定性。最后,我们在两个仿真基准和五个真实世界环境任务中进行了评估,实验结果验证了所提算法的有效性。

原文摘要 · Abstract (English)

The wide application of flow-matching methods has greatly promoted the development of robot imitation learning. However, these methods all face the problem of high inference time. To address this issue, researchers have proposed distillation methods and consistency methods, but the performance of these methods still struggles to compete with that of the original diffusion models and flow-matching models. In this article, we propose a one-step shortcut method with multi-step integration for robot imitation learning. To balance the inference speed and performance, we extend the multi-step consistency loss on the basis of the shortcut model, split the one-step loss into multi-step losses, and improve the performance of one-step inference. Secondly, to solve the problem of unstable optimization of the multi-step loss and the original flow-matching loss, we propose an adaptive gradient allocation method to enhance the stability of the learning process. Finally, we evaluate the proposed method in two simulation benchmarks and five real-world environment tasks. The experimental results verify the effectiveness of the proposed algorithm.

模仿学习加速推理机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。