arXiv:2606.24745cs.SDcs.AI2026-06

用声学特征对齐替代跳跃连接,提升语音增强的实时性与质量。

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

论文配图:Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement
图 1 · 摘自论文原文
  • 摒弃传统跳跃连接,通过潜在表示对齐引导编码器-解码器结构
  • 仅用5次函数求值即在VoiceBank-DEMAND上实现更优的PESQ与感知质量
  • 适合追求高效低延迟语音增强的系统集成者

生成模型如扩散模型和基于分数的方法在语音增强中表现优异,但其迭代采样过程限制了实时部署。流匹配通过常微分方程以极少函数求值即可将噪声语音映射至干净语音,是一种高效替代方案。本文提出一种无跳跃连接的编码器-解码器主干网络,基于潜在表示对齐(LRA)设计。该模型不依赖可能传递噪声相关低级特征的U-Net跳跃连接,而是通过冻结的Descript Audio Codec编码器-解码器(无量化)提取干净语音潜在特征,对齐瓶颈层与解码器表示。这种对齐监督促使紧凑的干净语音表示,同时保持高效的少步推理。在WSJ0-CHiME3与VoiceBank-DEMAND数据集上的实验表明,该方法显著提升了PESQ与感知质量,尤其在VoiceBank-DEMAND上效果突出,且仅需5次函数求值。

原文摘要 · Abstract (English)

Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but their iterative sampling process limits real-time deployment. Flow Matching offers an efficient alternative by transporting noisy speech toward clean speech through an ordinary differential equation with few function evaluations. In this work, we propose a skip-free encoder-decoder backbone for flow-matching speech enhancement, guided by Latent Representation Alignment (LRA). Instead of relying on U-Net skip connections, which may transfer noise-correlated low-level features to the decoder, the proposed model aligns its bottleneck and decoder representations with clean latent features extracted from a frozen Descript Audio Codec encoder-decoder without quantization. This codec-aligned supervision promotes compact clean-speech representations while preserving efficient few-step inference. Experiments on WSJ0-CHiME3 and VoiceBank-DEMAND show improved PESQ and perceptual quality, especially on VoiceBank-DEMAND, using only five function evaluations.

语音增强流匹配潜在对齐高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。