用音频生成全身动作与口型同步的真人视频
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
- 分阶段扩散变换器架构,先生成全身动态,再细化手脸细节
- 实现高保真视频,口型与音频同步,肢体动作自然连贯
- 适合影视动画、虚拟主播等需要精准语音驱动视频场景
尽管音频驱动视频生成取得进展,现有方法多聚焦面部动作,导致头部与身体运动不协调。本文提出AudCast,一种基于参考图像和音频生成全身人体视频的通用框架,采用级联扩散变换器(DiTs)范式:首先设计音频条件的全身人体DiT,直接驱动任意人体的生动手势动态;随后引入区域精修DiT,利用区域3D拟合作为桥梁重构信号,提升手部和面部细节。大量实验表明,该框架能生成高保真、时间连贯且包含精细面部与手部细节的音频驱动全身视频。资源见 https://guanjz20.github.io/projects/AudCast。
原文摘要 · Abstract (English)
Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。