用想象的动态规划让机器人完成复杂连续操作任务
LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

- 通过学习可复用技能与动态模型,实现高层规划
- 在多物体重排任务中成功率达78%,部分完成率提升23%
- 适合需要长时序动作组合的机器人控制研究
长时序人形机器人行走操控需要组合多种全身技能并做出可靠高层决策。现有方法通常将预训练技能与脚本化规划器、有限状态机或特定任务的无模型策略结合,限制了其处理复杂任务序列的能力。为此,我们提出LUCID——一种分层模型化强化学习框架,通过学习到的动态模型进行想象式推演来规划可复用技能。LUCID首先通过对抗性模仿训练结构化的隐变量条件低层策略,随后冻结该策略,联合学习高层策略与宏观动态世界模型。世界模型预测由隐变量决策引发的时序状态转移,使高层策略可通过想象式推演进行优化。我们在多个模拟的多物体重排场景中评估该框架。实验结果表明,相比先前基线方法,LUCID在完整任务成功率和部分完成率上均有提升,验证了其在复杂序列式人形操控任务中的有效性。
原文摘要 · Abstract (English)
Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often coordinate pretrained skills with scripted planners, finite-state machines or task-specific model-free policies, restricting their ability to handle complex task sequences. To address this limitation, we propose \textbf{LUCID}, a hierarchical model-based reinforcement learning framework that plans over reusable skills through imagined rollouts of a learned dynamics model. LUCID first trains a structured latent-conditioned low-level policy via adversarial imitation and then freezes it while jointly learning a high-level policy and macro-dynamics world model. The world model predicts the temporally extended state transitions induced by latent decisions, enabling high-level policy optimization through imagined rollouts. We evaluate our framework across various simulated multi-object rearrangement scenarios. Experimental results show that LUCID improves the full-task success and partial-completion rates compared to prior baseline methods, demonstrating its effectiveness in complex sequential loco-manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。