arXiv:2503.08007cs.ROcs.AI2025-03ICRA被引 27

用强化学习让四足机器人高效学多种技能,还能在真实环境落地。

MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

  • 多专家混合架构,低秩适配模块动态激活,支持多任务灵活切换。
  • 在六项技能上超越所有基线,分布外泛化能力显著提升。
  • 适合研究多任务机器人控制、强化学习与视觉语言模型融合的学者。

构建能在真实环境中流畅执行多种动作和任务的通用四足机器人仍是重大挑战。本文提出一种新型视觉-语言-动作(VLA)模型——混合机器人专家(MoRE),旨在利用强化学习对大规模VLA模型进行微调,并有效利用大量质量参差的数据。MoRE在密集的多模态大语言模型中集成多个低秩适应模块作为独立专家,形成稀疏激活的专家混合模型,从而实现对多样化下游任务的有效适应。此外,我们基于任务结构特性设计了强化学习训练目标,将模型训练为一个Q函数。通过自动采集的混合质量数据进行有效学习,显著提升了数据效率与模型性能。大量实验表明,MoRE在六种不同技能上均优于所有基线,并在分布外场景中展现出优越的泛化能力。我们在真实场景中进一步验证了该方法的可行性,证实其实际应用价值,为未来四足机器人多任务学习研究奠定了坚实基础。

原文摘要 · Abstract (English)

Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinforcement learning (RL) for fine-tuning large-scale VLA models with a large amount of mixed-quality data. MoRE integrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model. This design enables the model to effectively adapt to a wide array of downstream tasks. Moreover, we employ a reinforcement learning-based training objective to train our model as a Q-function after deeply exploring the structural properties of our tasks. Effective learning from automatically collected mixed-quality data enhances data efficiency and model performance. Extensive experiments demonstrate that MoRE outperforms all baselines across six different skills and exhibits superior generalization capabilities in out-of-distribution scenarios. We further validate our method in real-world scenarios, confirming the practicality of our approach and laying a solid foundation for future research on multi-task learning in quadruped robots.

四足机器人强化学习多模态模型技能泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。