从人脸视频生成高质量语音,分三步建模声音特征。
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
- 分内容、音色、语调三阶段逐步转化视频为语音特征。
- 生成语音与真实发音相似度高,优于现有方法。
- 适合语音合成、人像驱动语音研究者参考。
本研究旨在从无声说话人脸视频中生成高质量语音,即视频到语音合成。该任务面临无声视频与多维度语音之间显著的模态差异。为此,本文提出一种新系统,通过学习从视频到语音的分层表示来有效弥合这一差距。具体而言,通过三个连续阶段——内容、音色和语调建模,将无声视频逐步映射至声学特征空间。每个阶段均对齐唇部动作、面部身份和面部表情等视觉因素与对应的声学特征,确保转换流畅。此外,为从视觉表示生成真实且连贯的语音,采用流匹配模型,直接估计从简单先验分布到目标语音分布的轨迹。大量实验表明,该方法生成语音质量极佳,接近真实语音,显著优于现有方法。
原文摘要 · Abstract (English)
The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enhancing the quality of synthesized speech. This is achieved by learning of hierarchical representations from video to speech. Specifically, we gradually transform silent video into acoustic feature spaces through three sequential stages -- content, timbre, and prosody modeling. In each stage, we align visual factors -- lip movements, face identity, and facial expressions -- with corresponding acoustic counterparts to ensure the seamless transformation. Additionally, to generate realistic and coherent speech from the visual representations, we employ a flow matching model that estimates direct trajectories from a simple prior distribution to the target speech distribution. Extensive experiments demonstrate that our method achieves exceptional generation quality comparable to real utterances, outperforming existing methods by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。