arXiv:2512.13495cs.CV2025-12被引 4

用一张人脸图+文本音频生成高保真长时序数字人动画

Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation

  • 多模态驱动,融合图像、文本、音频生成连贯表情与口型
  • 在100万标注数据上训练,实现身份一致与精准对口型
  • 推理速度提升11.4倍,适合虚拟主播等实际场景

我们提出Soul框架,基于单张肖像图、文本提示和音频,生成语义连贯的高保真长时序数字人视频,实现精准口型同步、生动面部表情和强身份保持。构建包含100万条精细标注样本的Soul-1M数据集,涵盖头像、上半身、全身及多人场景,通过自动化标注流程缓解数据稀缺问题;并设计Soul-Bench用于公平评估音频/文本引导的动画方法。模型基于Wan2.2-5B骨干网络,集成音频注入层与多种训练策略,并采用阈值感知代码本替换以保障长期生成一致性。结合步数/CFG蒸馏与轻量级VAE优化推理效率,实现11.4倍加速且质量损失可忽略。大量实验表明,Soul在视频质量、图文对齐、身份保持和口型同步精度上显著优于现有开源与商用模型,在虚拟主播、影视制作等真实场景中具备广泛应用潜力。

原文摘要 · Abstract (English)

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed $\textbf{Soul}$, which generates semantically coherent videos from a single-frame portrait image, text prompts, and audio, achieving precise lip synchronization, vivid facial expressions, and robust identity preservation. We construct Soul-1M, containing 1 million finely annotated samples with a precise automated annotation pipeline (covering portrait, upper-body, full-body, and multi-person scenes) to mitigate data scarcity, and we carefully curate Soul-Bench for comprehensive and fair evaluation of audio-/text-guided animation methods. The model is built on the Wan2.2-5B backbone, integrating audio-injection layers and multiple training strategies together with threshold-aware codebook replacement to ensure long-term generation consistency. Meanwhile, step/CFG distillation and a lightweight VAE are used to optimize inference efficiency, achieving an 11.4$\times$ speedup with negligible quality loss. Extensive experiments show that Soul significantly outperforms current leading open-source and commercial models on video quality, video-text alignment, identity preservation, and lip-synchronization accuracy, demonstrating broad applicability in real-world scenarios such as virtual anchors and film production. Project page at https://zhangzjn.github.io/projects/Soul/

数字人多模态生成口型同步高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。