arXiv:2505.22024cs.SDcs.CV2025-05中稿 · Interspeech 2025被引 1

分离声学与语义特征,让无声口型视频生成更自然的语音。

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

  • 拆分语音生成路径:声学路径管语调,语义路径抓内容。
  • 在两个标准数据集上均实现更高可懂度和语音自然度。
  • 适合做语音重建、口型驱动语音生成的研究者参考。

唇读语音(L2S)合成旨在从视觉线索中重建语音,但因缺乏足够监督,难以准确捕捉语言内容、口音和语调。本文提出 RESOUND,一种新型 L2S 系统,能从无声说话人脸视频生成清晰且富有表现力的语音。基于声源-滤波器理论,方法包含两个组件:声学路径用于预测语调,语义路径用于提取语言特征。这种分离简化了学习过程,支持对两类表征独立优化。此外,通过将语音单元(speech units)——一种成熟的无监督语音表示技术——与梅尔频谱图一同引入波形生成,显著提升性能。该设计使 RESOUND 在保留内容与说话人身份的同时,生成具有丰富语调的语音。在两个标准 L2S 基准数据集上的实验验证了该方法的有效性,各项指标均有提升。

原文摘要 · Abstract (English)

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a novel L2S system that generates intelligible and expressive speech from silent talking face videos. Leveraging source-filter theory, our method involves two components: an acoustic path to predict prosody and a semantic path to extract linguistic features. This separation simplifies learning, allowing independent optimization of each representation. Additionally, we enhance performance by integrating speech units, a proven unsupervised speech representation technique, into waveform generation alongside mel-spectrograms. This allows RESOUND to synthesize prosodic speech while preserving content and speaker identity. Experiments conducted on two standard L2S benchmarks confirm the effectiveness of the proposed method across various metrics.

唇读语音语音生成多模态声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。