arXiv:2512.17253cs.CV2025-12被引 11

无需关键点或动作标签,直接从人类示范视频生成机器人执行视频

Mitty: Diffusion-based Human-to-Robot Video Generation

  • 基于扩散模型的Transformer架构,端到端完成人到机器人的视频生成
  • 在Human2Robot和EPIC-Kitchens数据集上达到当前最优性能,泛化能力出色
  • 适用于希望从人类示范中学习的机器人研究者,尤其关注视觉一致性

从人类示范视频中直接学习是实现可扩展、通用机器人学习的关键。然而,现有方法依赖关键点或轨迹等中间表示,引入信息丢失和累积误差,损害时序与视觉一致性。我们提出Mitty,一种基于预训练视频扩散模型的扩散Transformer,支持端到端的人类示范视频到机器人执行视频的上下文学习。演示视频被压缩为条件令牌,通过双向注意力与机器人去噪令牌融合。为缓解成对数据稀缺问题,我们还构建了自动化合成流程,利用大规模第一视角数据集生成高质量的人机配对视频。在Human2Robot和EPIC-Kitchens上的实验表明,Mitty实现了当前最佳表现,具备强泛化能力,并为从人类观察中实现可扩展机器人学习提供了新见解。

原文摘要 · Abstract (English)

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss and cumulative errors that harm temporal and visual consistency. We present Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation. Built on a pretrained video diffusion model, Mitty leverages strong visual-temporal priors to translate human demonstrations into robot-execution videos without action labels or intermediate abstractions. Demonstration videos are compressed into condition tokens and fused with robot denoising tokens through bidirectional attention during diffusion. To mitigate paired-data scarcity, we also develop an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets. Experiments on Human2Robot and EPIC-Kitchens show that Mitty delivers state-of-the-art results, strong generalization to unseen environments, and new insights for scalable robot learning from human observations.

视频生成扩散模型机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。