arXiv:2605.15336cs.ROcs.AI2026-05被引 6

HoloMotion-1可零样本追踪任意动作,从真实视频学运动,直接控制真机器人。

HoloMotion-1 Technical Report

论文配图:HoloMotion-1 Technical Report
图 1 · 摘自论文原文
  • 用真实视频重建动作+人工采集数据混合训练,提升动作多样性
  • 在多个未见基准上显著提升追踪精度,实测可直接部署于真实机器人
  • 采用稀疏专家变换器与缓存推理,实现低延迟实时控制

本报告介绍HoloMotion-1,一种面向零样本全身动作追踪的人形运动基础模型。其核心创新在于利用大规模混合动作数据集进行控制策略训练:来自野外视频的动作重建数据提供主要的动作多样性,而精心筛选的动捕数据和自研动作数据提供高保真监督与部署适配覆盖。该数据范式使模型突破传统仅依赖动捕数据的局限,暴露于更广泛的动作行为、采集条件与风格中。学习此类异构数据带来新挑战,包括重建噪声、源域差异、质量不均及大行为变化下的时序建模。为此,HoloMotion-1融合大容量时序建模、基于键值缓存的稀疏激活专家混合变压器以实现实时控制,并采用序列级训练策略提升长序列学习效率。在多个未见动作基准上的大量实验表明,HoloMotion-1能鲁棒泛化至多样化动作类型与采集条件,追踪精度显著优于先前方法,并可直接迁移至真实人形机器人,无需任务微调。

原文摘要 · Abstract (English)

In this report, we present HoloMotion-1, a humanoid motion foundation model for zero-shot whole-body motion tracking. A key innovation of HoloMotion-1 is to scale control-policy training with a large-scale hybrid motion corpus, where video-reconstructed motions from in-the-wild videos provide the dominant source of motion diversity, while curated motion-capture and in-house motion data provide higher-fidelity supervision and deployment-oriented coverage. This data regime enables HoloMotion-1 to move beyond conventional MoCap-only training and exposes the policy to substantially broader behaviors, capture conditions, and motion styles. Learning from such heterogeneous data introduces new challenges, including reconstruction noise, source-domain mismatch, uneven motion quality, and the need for temporal modeling under large behavioral variation. To address these challenges, HoloMotion-1 integrates large-capacity temporal modeling, a sparsely activated Mixture-of-Experts Transformer with KV-cache inference for real-time control, and a sequence-level training strategy that improves learning efficiency on extended motion sequences. Extensive experiments on multiple unseen motion benchmarks show that HoloMotion-1 generalizes robustly across diverse motion types and capture conditions, significantly improves tracking accuracy over prior methods, and transfers directly to a real humanoid robot without task-specific fine-tuning.

动作生成零样本人形机器人视频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。