arXiv:2503.17626cs.ROcs.AI2025-03被引 2

用统一潜空间策略让不同机器人快速学会走路,省时又通用。

Transferable Latent-to-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged Robots

  • 先在潜空间预训练通用行走策略,再微调特定编码器解码器。
  • 实测新机器人仅需少量数据即可高效适应新任务和新体型。
  • 适合需要快速部署的多形态足式机器人应用场景。

强化学习在获取机器人技能方面表现优异,但每个新技能仍需大量数据训练。采用预训练-微调范式可高效适配新机器人与新任务。受已有知识能加速同机器人学新任务、或让新机器人快速掌握已训练任务的启发,我们提出一种潜空间训练框架:预训练一个可迁移的潜变量到潜变量的行走策略,搭配多样化的任务专用观测编码器与动作解码器。该策略在潜空间中处理编码后的潜变量观测,生成潜变量动作并解码输出,具备学习通用抽象运动技能的潜力。为保留决策与控制的关键信息,引入扩散恢复模块,在预训练阶段最小化信息重建损失。微调阶段,固定预训练的潜变量行走策略,仅优化轻量级的任务专用编码器与解码器,实现高效适应。方法使机器人能够利用自身跨任务经验,以及其它形态各异机器人的经验,加速适应过程。通过大量仿真与真实实验验证,预训练的潜变量到潜变量行走策略能有效泛化至新机器人实体与新任务,显著提升效率。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has demonstrated remarkable capability in acquiring robot skills, but learning each new skill still requires substantial data collection for training. The pretrain-and-finetune paradigm offers a promising approach for efficiently adapting to new robot entities and tasks. Inspired by the idea that acquired knowledge can accelerate learning new tasks with the same robot and help a new robot master a trained task, we propose a latent training framework where a transferable latent-to-latent locomotion policy is pretrained alongside diverse task-specific observation encoders and action decoders. This policy in latent space processes encoded latent observations to generate latent actions to be decoded, with the potential to learn general abstract motion skills. To retain essential information for decision-making and control, we introduce a diffusion recovery module that minimizes information reconstruction loss during pretrain stage. During fine-tune stage, the pretrained latent-to-latent locomotion policy remains fixed, while only the lightweight task-specific encoder and decoder are optimized for efficient adaptation. Our method allows a robot to leverage its own prior experience across different tasks as well as the experience of other morphologically diverse robots to accelerate adaptation. We validate our approach through extensive simulations and real-world experiments, demonstrating that the pretrained latent-to-latent locomotion policy effectively generalizes to new robot entities and tasks with improved efficiency.

机器人控制强化学习迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。