无需分拆上下身,用GPT模型实现多形态机器人全身动作精准追踪
VENOM: Versatile Embodied Network for Omni-bodied Motion tracking

- 基于GPT架构的统一模型,直接端到端追踪全身动作
- 在多形态机器人上实现稳定追踪,性能超越纯监督学习的MLP模型
- 无需奖励反馈,接近强化学习专家水平,适合仿真环境动作迁移
在仅依赖示范数据的情况下,实现跨多种人形机器人平台的专家级全身动作追踪仍是一个极具挑战且研究较少的问题。现有跨形态动作追踪策略通常将控制问题解耦为上下身分别控制。本文提出VENOM,一种面向仿真环境中多形态人形机器人的跨形态全身动作追踪模型。VENOM是一种基于GPT的运动追踪器,通过在多个形态人形机器人数据上训练,可无需分拆上下身直接追踪完整身体动作。我们构建了名为VENOM的数据集,包含状态、动作和奖励信息,并在此数据集上训练VENOM及基线模型。实验表明,相比仅使用监督学习的多形态数据训练的MLP模型,VENOM在不同人形机器人上表现出更优且稳定的追踪性能;同时,尽管未使用奖励反馈,其追踪能力仍可媲美采用非对称演员-评论家强化学习训练的专家模型。
原文摘要 · Abstract (English)
Achieving expert-level expressive full-body motion tracking across multiple humanoids solely from demonstration data remains a challenging and relatively an underexplored problem in humanoid robot learning. Cross-embodiment motion tracking policies are mostly trained by decoupling the control problem into upper and lower body control. This work proposes VENOM, a cross-embodiment full-body motion tracking model for humanoids in simulation. VENOM is a GPT-based motion tracker trained on multiple humanoid data that can track the entire body without the requirement to split into upper and lower body control. We curate a multi-humanoid motion tracking dataset called the VENOM dataset that contains states, actions, and rewards and train VENOM and the baselines on this dataset. In this letter, we evaluate VENOM's performance against baselines and show that we can achieve a stable motion tracker across different humanoids more capable than an MLP trained on multiple humanoid data with supervised learning alone, and also show that despite lack of reward feedback, VENOM closely matches the tracking capability of experts that were trained using asymmetric-actor critic reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。