用本体感觉增强表示学习,让机器人更省数据地学会复杂动作
PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations
- 利用本体感觉与特权状态的互补性,自适应学出紧凑任务相关表示
- 在真实机器人上实现更快收敛,样本效率提升显著,性能超越基线方法
- 适合追求高效机器人学习的研究者和工程团队
实现高效且鲁棒的全身控制(WBC)对于使类人机器人在动态环境中完成复杂任务至关重要。尽管强化学习(RL)在此领域取得成功,但其样本效率低仍是主要挑战,原因在于类人机器人的复杂动力学和部分可观测性。为此,我们提出PvP——一种本体感觉特权对比学习框架,利用本体感觉与特权状态之间的内在互补性。PvP无需手工设计数据增强即可学习紧凑且任务相关的潜在表示,从而实现更快、更稳定的策略学习。为支持系统性评估,我们开发了SRL4Humanoid,首个统一且模块化的框架,提供代表性状态表示学习(SRL)方法在类人机器人学习中的高质量实现。在LimX Oli机器人上进行的速度追踪和运动模仿任务的大量实验表明,PvP相比基线SRL方法显著提升了样本效率和最终性能。本研究进一步为将SRL与RL结合用于类人机器人WBC提供了实用洞见,为数据高效的类人机器人学习提供重要指导。
原文摘要 · Abstract (English)
Achieving efficient and robust whole-body control (WBC) is essential for enabling humanoid robots to perform complex tasks in dynamic environments. Despite the success of reinforcement learning (RL) in this domain, its sample inefficiency remains a significant challenge due to the intricate dynamics and partial observability of humanoid robots. To address this limitation, we propose PvP, a Proprioceptive-Privileged contrastive learning framework that leverages the intrinsic complementarity between proprioceptive and privileged states. PvP learns compact and task-relevant latent representations without requiring hand-crafted data augmentations, enabling faster and more stable policy learning. To support systematic evaluation, we develop SRL4Humanoid, the first unified and modular framework that provides high-quality implementations of representative state representation learning (SRL) methods for humanoid robot learning. Extensive experiments on the LimX Oli robot across velocity tracking and motion imitation tasks demonstrate that PvP significantly improves sample efficiency and final performance compared to baseline SRL methods. Our study further provides practical insights into integrating SRL with RL for humanoid WBC, offering valuable guidance for data-efficient humanoid robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。