arXiv:2607.26460cs.RO2026-07

用隐空间强化学习提升机器人移动操作的运动规划能力

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

论文配图:RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning
图 1 · 摘自论文原文
  • 先用专家数据预训练流模型,再通过隐空间优化生成更优动作
  • 在基准测试中任务成功率、避障和轨迹质量显著优于纯模仿方法
  • 适合需要高效且高质量运动规划的移动机器人研究者

移动操作需要生成满足目标达成、避障、底盘运动学约束、机械臂关节极限和轨迹平滑性的全身动作序列。基于流的生成策略能从专家示范中高效学习多模态且时间一致的动作先验,但仅依赖模仿学习无法超越示范分布的质量。我们提出RLMM-Flow,一种结合专家流策略预训练与隐空间强化学习后训练的框架。该框架首先从专家示范中学习一个捕捉多模态全身运动先验的流策略;随后冻结预训练流策略,由一个隐空间引导网络将初始噪声导向高价值动作片段。为稳定高维隐空间优化,我们先预热动作空间评价器,再联合训练隐空间评价器与隐空间策略,并引入粗到细的隐空间引导机制,逐步从共享时域隐表示扩展至全维度残差表示。在移动操作运动规划基准上实验表明,相较于仅模仿的流策略及现有强化学习后训练基线,RLMM-Flow显著提升了任务成功率、避障能力和轨迹质量,同时保持了流模型的快速推理特性。

原文摘要 · Abstract (English)

Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, manipulator joint limits, and trajectory smoothness. Flow-based generative policies provide an efficient paradigm for learning multimodal and temporally consistent motion priors from expert demonstrations, but imitation-only training cannot improve policy quality beyond the demonstration distribution. We propose RLMM-Flow, a flow-based mobile manipulation framework that combines expert flow-policy pretraining with latent-space reinforcement learning post-training. The framework first learns a flow policy that captures a multimodal whole-body motion prior from expert demonstrations. The pretrained flow policy is then frozen, while a latent steering network steers its initial noise toward higher-value action chunks. To stabilize high-dimensional latent optimization, we warm up an action-space critic before jointly training the latent critic and latent actor, and introduce coarse-to-fine latent steering that progressively expands control from a horizon-shared latent representation to a full-dimensional residual representation. Experiments on mobile manipulation motion-planning benchmarks show that RLMM-Flow substantially improves task success, collision avoidance, and trajectory quality over imitation-only flow policies and existing reinforcement learning post-training baselines, while preserving fast flow-based inference.

机器人强化学习运动规划流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。