用第一视角人类示范训练机器人全身动作生成模型
EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

- 三流DiT联合建模身体动作、第一视角视觉与文本
- 单个检查点实现动作重建、生成与预测,优于基线方法
- 适合需要交互式人形机器人控制的开发场景
人形机器人需具备适应场景、任务和用户意图的全身动作能力。运动追踪可复现特定轨迹,视觉-语言-动作系统提供语义接口,但均无法为广泛全身行为提供可扩展且交互式的先验知识。本文提出EgoPriMo(第一视角人形动作先验),从第一视角人类示范中学习此类先验。给定第一视角观测与文本提示,EgoPriMo可重建、生成并预测基于SMPL的全身动作。语言作为高层控制信号而非完整动作说明。核心是三流DiT,联合建模身体动力学、第一视角视觉上下文与文本;任务条件掩码将不同任务和缺失模态数据引导至同一检查点。在Nymeria与EgoExo4D上的实验表明,单一检查点在第一视角动作生成上优于UniEgoMotion,同时支持重建与预测;生成的SMPL动作可由Unitree人形控制器执行。结果表明,从可扩展的第一视角观测到通用且交互式的人形动作先验,已具可行性。
原文摘要 · Abstract (English)
Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic interfaces, but neither offers a scalable and interactive prior for broad full-body behavior. We introduce EgoPriMo (Egocentric Motion Prior for Humanoid Robots), a unified framework that learns such priors from egocentric human demonstrations. Given egocentric observations and a text prompt, EgoPriMo reconstructs, generates, and forecasts SMPL-based full-body motion. Language is used as a high-level control signal rather than a complete motion specification. At the core of EgoPriMo is a Triple-stream DiT that jointly models body dynamics, egocentric visual context, and text; task-conditioning masks route different tasks and missing-modality data through the same checkpoint. Experiments on Nymeria and EgoExo4D show that one checkpoint improves egocentric motion generation over UniEgoMotion while supporting reconstruction and forecasting; the generated SMPL motions can also be executed by a Unitree humanoid controller. These results indicate a practical path from scalable egocentric observations to generalizable and interactive humanoid motion priors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。