让真人视频生成更真实、灵活,支持多种驱动方式和姿态
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

- 用混合运动条件训练扩散变换器,提升数据利用效率
- 生成视频逼真度高,支持口型/歌声/肢体动作/交互等复杂场景
- 适合需要多模态输入、高自由度真人动画的开发者和研究者
端到端的人体动画近年来取得显著进展,如音频驱动的说话人生成。然而,现有方法难以像大规模视频生成模型一样扩展,限制了实际应用。本文提出OmniHuman,一种基于扩散变压器的框架,通过在训练阶段融合运动相关条件来扩大数据规模。为此,我们设计了两种训练原则,以及相应的模型架构与推理策略,使OmniHuman能充分挖掘数据驱动的运动生成能力,最终实现高度真实的真人视频生成。更重要的是,OmniHuman支持多种人物画面(面部特写、半身、全身)、支持说话与歌唱、可处理人-物交互及复杂身体姿态,并兼容不同图像风格。相比现有端到端音频驱动方法,OmniHuman不仅生成更真实视频,还具备更强输入灵活性,支持音频驱动、视频驱动及组合驱动信号。视频样例见ttfamily项目页(https://omnihuman-lab.github.io)。
原文摘要 · Abstract (English)
End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related conditions into the training phase. To this end, we introduce two training principles for these mixed conditions, along with the corresponding model architecture and inference strategy. These designs enable OmniHuman to fully leverage data-driven motion generation, ultimately achieving highly realistic human video generation. More importantly, OmniHuman supports various portrait contents (face close-up, portrait, half-body, full-body), supports both talking and singing, handles human-object interactions and challenging body poses, and accommodates different image styles. Compared to existing end-to-end audio-driven methods, OmniHuman not only produces more realistic videos, but also offers greater flexibility in inputs. It also supports multiple driving modalities (audio-driven, video-driven and combined driving signals). Video samples are provided on the ttfamily project page (https://omnihuman-lab.github.io)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。