arXiv:2606.03985cs.ROcs.AI2026-06中稿 · CVPR被引 6

用百亿帧数据训练的生成模型,零样本追踪复杂人体动作

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

论文配图:Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
图 1 · 摘自论文原文
  • 基于因果注意力的Transformer,用20亿帧动捕数据预训练
  • 零样本泛化到未见动作与控制任务,同时稳定追踪高动态行为
  • 适合做通用人体运动生成与控制,尤其关注数据规模效应

我们提出Humanoid-GPT,一种基于因果注意力机制的GPT风格Transformer,其在百亿级动作语料上进行预训练,实现全身运动控制。与以往受数据稀缺和敏捷性-泛化权衡限制的浅层MLP跟踪器不同,Humanoid-GPT在包含所有主要动捕数据集及大规模自研记录的20亿帧重定向语料上进行预训练。通过同时扩展数据量与模型容量,该模型仅需单一生成式Transformer即可追踪高度动态的行为,并实现前所未有的零样本泛化能力,覆盖未见动作与控制任务。大量实验与缩放分析表明,该模型确立了新性能基准,在零样本泛化与复杂动态运动追踪方面均表现卓越。

原文摘要 · Abstract (English)

We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus that unifies all major mocap datasets with large-scale in-house recordings. Scaling both data and model capacity yields a single generative Transformer that tracks highly dynamic behaviors while achieving unprecedented zero-shot generalization to unseen motions and control tasks. Extensive experiments and scaling analyses show that our model establishes a new performance frontier, demonstrating robust zero-shot generalization to unseen tasks while simultaneously tracking highly dynamic and complex motions.

动作生成零样本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。