arXiv:2603.13707cs.ROcs.AI2026-03被引 3

用强化学习微调扩散策略,让机器人更稳地完成复杂动作任务。

REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning

  • 分层优化扩散规划器与强化学习控制器,协同提升动作精度。
  • 仿真中任务成功率超90%,且在未见过的场景仍表现稳定。
  • 适合需要高可靠性的人形机器人运动操控研究者使用。

人形机器人执行运动-操作任务需在复杂动态环境下实现协调的任务空间规划与稳定的动作跟踪。尽管扩散策略(DP)从示范中学习具有潜力,但在人形机器人部署时面临关键挑战:离线训练的运动规划器与运动控制模块解耦,导致命令跟踪不佳,加剧分布偏移并引发任务失败。常规的数据规模扩展方法对高维人形系统成本过高。为此,我们提出REFINE-DP(REinforcement learning FINE-tuning of Diffusion Policy),一种分层框架,联合优化扩散策略规划器与基于强化学习的运动控制模块。通过基于PPO的扩散策略梯度对规划器进行微调以提升任务成功率,同时同步更新控制器以准确跟踪规划器演化出的命令分布,缓解分布不匹配问题从而改善运动质量。我们在人形机器人上验证了该方法在门洞穿越和长时程物体运输等任务中的表现,结果显示在仿真中成功率超过90%,即使面对预训练数据中未出现的分布外情况也保持稳健,并可在真实世界中执行而无需特权状态信息。所提方法显著优于预训练的扩散策略基线,证明强化学习微调对可靠人形运动-操作至关重要。

原文摘要 · Abstract (English)

Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon tasks. While diffusion policies (DPs) show promise for learning from demonstrations, deploying them on humanoids poses critical challenges: the motion planner trained offline is decoupled from the loco-manipulation controller, leading to poor command tracking, compounding distribution shift, and task failures. The common approach of scaling demonstration data is prohibitively expensive for high-dimensional humanoid systems. To address this challenge, we present REFINE-DP (REinforcement learning FINE-tuning of Diffusion Policy), a hierarchical framework that jointly optimizes a DP motion planner and an RL-based loco-manipulation controller. The DP is fine-tuned via a PPO-based diffusion policy gradient to improve task success rate, while the controller is simultaneously updated to accurately track the planner's evolving command distribution, reducing the distributional mismatch that degrades motion quality. We validate REFINE-DP on a humanoid robot performing loco-manipulation tasks, including door traversal and long-horizon object transport. REFINE-DP achieves an over 90% success rate in simulation, even in out-of-distribution cases not seen in the pre-training data, and enables real-world execution without privileged state information. Our proposed method substantially outperforms pre-trained DP baselines and demonstrates that RL fine-tuning is key to reliable humanoid loco-manipulation. https://refine-dp.github.io/REFINE-DP/

人形机器人扩散模型强化学习运动规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。