arXiv:2604.10165cs.RO2026-04

混合强化与模仿学习,让机器人高效完成复杂抓取任务。

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

  • 根据动作方差动态切换模仿与强化学习专家
  • 真实场景下平均成功率97.5%,2~5小时完成微调
  • 大幅减少人工干预,适合工业级长时序操作

强化学习(RL)和模仿学习(IL)是机械臂策略获取的标准框架。虽然IL能高效生成策略,但存在误差累积和分布偏移问题;而RL虽可自主探索,却常因样本效率低、试错成本高而受限。针对复杂任务,本文提出混合强化与模仿学习专家(MoRI),通过动作方差动态切换IL与RL专家,以应对粗略运动与精细操作。系统采用离线预训练+在线微调策略加速收敛,并对RL组件施加基于IL的正则化以保障探索安全、减少人工干预。在四个真实世界复杂任务上的评估显示,MoRI在2至5小时内实现平均97.5%的成功率。相比基准RL算法,人类干预降低85.8%,收敛时间缩短21%,展现出强大的机器人操作能力。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient policy derivation, it suffers from compounding errors and distribution shift. Conversely, RL facilitates autonomous exploration but is frequently hindered by low sample efficiency and the high cost of trial and error. Since existing hybrid methods often struggle with complex tasks, we introduce Mixture of RL and IL Experts (MoRI). This system dynamically switches between IL and RL experts based on the variance of expert actions to handle coarse movements and fine-grained manipulations. MoRI employs an offline pre-training stage followed by online fine-tuning to accelerate convergence. To maintain exploration safety and minimize human intervention, the system applies IL-based regularization to the RL component. Evaluation across four complex real-world tasks shows that MoRI achieves an average success rate of 97.5% within 2 to 5 hours of fine-tuning. Compared to baseline RL algorithms, MoRI reduces human intervention by 85.8% and shortens convergence time by 21%, demonstrating its capability in robotic manipulation.

机器人操控强化学习模仿学习混合策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。