arXiv:2410.23836cs.CV2024-10TPAMI被引 4

用语音生成逼真3D人像视频,支持连续视角切换

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

  • 用大语言模型先验增强音频到动作的映射质量
  • 融合视图与区域引导的专家混合机制提升画面稳定性
  • 适合虚拟主播、数字人开发等需要自然肢体表达的场景

本文提出 Stereo-Talker,一种单次输入的音频驱动3D人像视频生成系统,可生成唇音同步精准、肢体表情丰富、时序一致且逼真的视频,并支持连续视角控制。该系统采用两阶段流程:第一阶段将音频映射为高保真上半身动作序列,结合大语言模型(LLM)先验与对齐文本语义的音频特征,利用LLM的跨模态泛化能力提升动作质量;第二阶段改进基于扩散的视频生成模型,引入先验引导的专家混合(MoE)机制——视图引导的MoE关注视角特异性属性,掩码引导的MoE增强局部渲染稳定性。同时设计掩码预测模块,从动作数据中生成人体掩码,提升掩码精度与推理期间的引导效果。构建包含2,203个身份的综合性人体视频数据集,涵盖多样姿态与详细标注,支持广泛泛化。代码、数据及预训练模型将公开用于研究。

原文摘要 · Abstract (English)

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and continuous viewpoint control. The process follows a two-stage approach. In the first stage, the system maps audio input to high-fidelity motion sequences, encompassing upper-body gestures and facial expressions. To enrich motion diversity and authenticity, large language model (LLM) priors are integrated with text-aligned semantic audio features, leveraging LLMs' cross-modal generalization power to enhance motion quality. In the second stage, we improve diffusion-based video generation models by incorporating a prior-guided Mixture-of-Experts (MoE) mechanism: a view-guided MoE focuses on view-specific attributes, while a mask-guided MoE enhances region-based rendering stability. Additionally, a mask prediction module is devised to derive human masks from motion data, enhancing the stability and accuracy of masks and enabling mask guiding during inference. We also introduce a comprehensive human video dataset with 2,203 identities, covering diverse body gestures and detailed annotations, facilitating broad generalization. The code, data, and pre-trained models will be released for research purposes.

3D生成音频驱动扩散模型数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。