无需配对数据即可跨机器人形态生成视频,支持新机器人快速适配。
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

- 分离运动与形态因素,用无配对数据训练轻量适配器。
- 在合成和真实数据上均实现高运动保真度与形态一致性。
- 适合需快速部署到新机器人的场景,如机器人训练数据生成。
跨形态视频生成旨在将动作从一种人形体态(如人到机器人、机器人到机器人)迁移,以支持具身智能的可扩展数据生成。主要挑战在于运动动态部分可迁移,而外观和形态则具有体态特异性。现有方法常混淆这些因素,且多数需为每个目标体态提供成对数据,限制了对新机器人的可扩展性。本文提出OmniHumanoid框架,将可迁移的运动学习与体态特异性适应分离。该方法从跨多种体态、场景和视角的运动对齐成对视频中学习共享运动迁移模型,通过仅使用无配对视频的轻量体态适配器实现对新体态的适应。为减少运动迁移与体态适应间的干扰,引入分支隔离注意力机制,分离运动条件与体态特异性调制。此外,构建了一个合成跨体态数据集,包含在多样化人形资产、场景和视角下渲染的运动对齐成对视频。在合成与真实世界基准测试中,OmniHumanoid在保持强运动保真度与体态一致性的同时,实现了无需重训练共享运动模型的可扩展新体态适应。
原文摘要 · Abstract (English)
Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is that motion dynamics are partly transferable across embodiments, whereas appearance and morphology remain embodiment-specific. Existing approaches often entangle these factors, and many require paired data for every target embodiment, which limits scalability to new robots. We present OmniHumanoid, a framework that factorizes transferable motion learning and embodiment-specific adaptation. Our method learns a shared motion transfer model from motion-aligned paired videos spanning multiple embodiments, while adapting to a new embodiment using only unpaired videos through lightweight embodiment-specific adapters. To reduce interference between motion transfer and embodiment adaptation, we further introduce a branch-isolated attention design that separates motion conditioning from embodiment-specific modulation. In addition, we construct a synthetic cross-embodiment dataset with motion-aligned paired videos rendered across diverse humanoid assets, scenes, and viewpoints. Experiments on both synthetic and real-world benchmarks show that OmniHumanoid achieves strong motion fidelity and embodiment consistency, while enabling scalable adaptation to unseen humanoid embodiments without retraining the shared motion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。