arXiv:2602.11477eess.AScs.CE2026-02AAAI

用分层潜空间扩散模型直接生成语音,提升唇语转语音质量。

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

论文配图:SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis
图 1 · 摘自论文原文
  • 构建分层子空间潜变量扩散模型,直接映射唇动到音频编码潜空间。
  • 通过扩散卷积块增强子空间间交互,实现高质量语音生成。
  • 融合语言模型与语义损失,更适合语音合成任务,适合语音生成研究者。

尽管近年来唇语转语音(L2S)取得了显著进展,但当前最先进的方法通常依赖于中间表示,如梅尔频谱图或离散自监督学习(SSL)标记。潜在扩散模型(LDM)在该任务中的潜力尚未充分探索。本文提出SLD-L2S,一种基于分层子空间潜扩散模型的新型L2S框架。该方法旨在将视觉唇动直接映射到预训练神经音频编码器的连续潜空间,从而避免传统中间表示带来的信息损失。其核心是通过子空间分解模块启动的多并行子空间处理架构。为高效增强子空间内及子空间间的交互,设计了扩散卷积块(DiCB)作为网络主干。此外,采用重参数化流匹配技术直接生成目标潜向量,实现训练中对语音语言模型(SLM)和语义损失的合理引入,超越传统流匹配目标,提升合成语音质量。实验表明,SLD-L2S在多个基准数据集上达到最先进水平,在客观与主观评估中均优于现有方法。

原文摘要 · Abstract (English)

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in this task remains largely unexplored. In this paper, we introduce SLD-L2S, a novel L2S framework built upon a hierarchical subspace latent diffusion model. Our method aims to directly map visual lip movements to the continuous latent space of a pre-trained neural audio codec, thereby avoiding the information loss inherent in traditional intermediate representations. The core of our method is a hierarchical architecture that processes visual representations through multiple parallel subspaces, initiated by a subspace decomposition module. To efficiently enhance interactions within and between these subspaces, we design the diffusion convolution block (DiCB) as our network backbone. Furthermore, we employ a reparameterized flow matching technique to directly generate the target latent vectors. This enables a principled inclusion of speech language model (SLM) and semantic losses during training, moving beyond conventional flow matching objectives and improving synthesized speech quality. Our experiments show that SLD-L2S achieves state-of-the-art generation quality on multiple benchmark datasets, surpassing existing methods in both objective and subjective evaluations.

唇语转语音扩散模型潜空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。