用音频驱动生成电影级角色动画,表现更自然真实。
Wan-S2V: Audio-Driven Cinematic Video Generation
- 基于Wan模型构建音频驱动动画系统,提升角色表现力。
- 在影视级动画中显著优于Hunyuan-Avatar和Omnihuman等模型。
- 适用于长视频生成与精准口型同步编辑,应用灵活。
当前最先进的音频驱动角色动画方法在语音和歌唱场景中表现良好,但在复杂影视制作中仍显不足,难以实现细腻的角色互动、真实的肢体动作和动态镜头。为解决这一长期挑战,我们提出名为Wan-S2V的音频驱动模型,基于Wan架构,在电影级语境下显著提升了角色表现力与真实性。通过大量实验,对比Hunyuan-Avatar与Omnihuman等前沿模型,结果一致表明本方法性能更优。此外,我们还探索了该方法在长视频生成与精确口型同步编辑中的多样性应用。
原文摘要 · Abstract (English)
Current state-of-the-art (SOTA) methods for audio-driven character animation demonstrate promising performance for scenarios primarily involving speech and singing. However, they often fall short in more complex film and television productions, which demand sophisticated elements such as nuanced character interactions, realistic body movements, and dynamic camera work. To address this long-standing challenge of achieving film-level character animation, we propose an audio-driven model, which we refere to as Wan-S2V, built upon Wan. Our model achieves significantly enhanced expressiveness and fidelity in cinematic contexts compared to existing approaches. We conducted extensive experiments, benchmarking our method against cutting-edge models such as Hunyuan-Avatar and Omnihuman. The experimental results consistently demonstrate that our approach significantly outperforms these existing solutions. Additionally, we explore the versatility of our method through its applications in long-form video generation and precise video lip-sync editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。