arXiv:2409.13268cs.CV2024-09被引 4

适配中文语音生成,提升口型与表情同步效果

JoyHallo: Digital human model for Mandarin

  • 用中文wav2vec2提取音频特征,适配普通话
  • 半解耦结构提升唇形表情姿态关联建模效率
  • 支持中英双语生成,适合医疗等专业场景

在语音驱动视频生成中,制作普通话视频面临显著挑战。收集全面的普通话数据集困难,且普通话复杂的口型运动比英语更难建模。本研究从京东健康国际公司员工采集了29小时普通话语音视频,构建了jdh-Hallo数据集,涵盖不同年龄和语调,包含日常对话及专业医疗话题。为适配JoyHallo模型生成普通话内容,采用中文wav2vec2进行音频特征嵌入,并提出半解耦结构,以捕捉唇形、表情与姿态之间的多特征关联。该设计不仅提高信息利用效率,还将推理速度提升14.3%。值得注意的是,JoyHallo仍保持生成英文视频的强能力,展现出优异的跨语言生成性能。代码与模型已开源:https://jdh-algo.github.io/JoyHallo。

原文摘要 · Abstract (English)

In audio-driven video generation, creating Mandarin videos presents significant challenges. Collecting comprehensive Mandarin datasets is difficult, and the complex lip movements in Mandarin further complicate model training compared to English. In this study, we collected 29 hours of Mandarin speech video from JD Health International Inc. employees, resulting in the jdh-Hallo dataset. This dataset includes a diverse range of ages and speaking styles, encompassing both conversational and specialized medical topics. To adapt the JoyHallo model for Mandarin, we employed the Chinese wav2vec2 model for audio feature embedding. A semi-decoupled structure is proposed to capture inter-feature relationships among lip, expression, and pose features. This integration not only improves information utilization efficiency but also accelerates inference speed by 14.3%. Notably, JoyHallo maintains its strong ability to generate English videos, demonstrating excellent cross-language generation capabilities. The code and models are available at https://jdh-algo.github.io/JoyHallo.

数字人语音生成中文建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。