从海量人类演示中学习动作意图先验,提升机器人操作的自然性和鲁棒性。
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation

- 构建分层视觉-语言-动作框架,分离人类动作中的意图与身体特征
- 在220万条动作-语言数据上训练,显著提升动作合理性与分布外适应能力
- 适合做具身智能、人形机器人操控的研究者参考
人类视频蕴含丰富的操作先验,但直接用于机器人学习存在挑战:原始观察混淆了场景理解、人体运动和个体身体特征。本文提出MoT-HRA,一种分层视觉-语言-动作框架,从大规模人类演示中学习人类意图先验。首先构建了包含220万条动作片段的HA-2.2M数据集,通过以手为中心的筛选、空间重建、时间分割和语言对齐,从异构人类视频中重构而成。在此基础上,MoT-HRA将操作分解为三个耦合专家:视觉-语言专家预测无身体特性的3D轨迹,意图专家将手部运动建模为类似MANO的隐式人类运动先验,精细专家将意图感知表征映射为机器人动作块。共享注意力主干与只读键值传递机制,使下游控制可利用人类先验,同时避免干扰上游表示。在手部运动生成、模拟操作及真实机器人任务上的实验表明,MoT-HRA显著提升了动作合理性与分布外鲁棒性。
原文摘要 · Abstract (English)
Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical vision-language-action framework that learns human-intention priors from large-scale human demonstrations. We first curate HA-2.2M, a 2.2M-episode action-language dataset reconstructed from heterogeneous human videos through hand-centric filtering, spatial reconstruction, temporal segmentation, and language alignment. On top of this dataset, MoT-HRA factorizes manipulation into three coupled experts: a vision-language expert predicts an embodiment-agnostic 3D trajectory, an intention expert models MANO-style hand motion as a latent human-motion prior, and a fine expert maps the intention-aware representation to robot action chunks. A shared-attention trunk and read-only key-value transfer allow downstream control to use human priors while limiting interference with upstream representations. Experiments on hand motion generation, simulated manipulation, and real-world robot tasks show that MoT-HRA improves motion plausibility and robust control under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。