arXiv:2606.17806eess.AS2026-06中稿 · Interspeech 2026

在语音表示空间中用流匹配提升语音质量,仅需4步采样即高效实现高保真还原。

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement

  • 直接在自监督语音模型的潜在空间建模音素条件下的清洁声学特征分布
  • 在语音增强任务中显著提升感知质量与可懂度,且仅需4次采样步骤
  • 适合追求高效推理的语音处理应用,如实时通话或边缘设备部署

流匹配(Flow Matching, FM)能够实现高保真生成,而自监督学习(Self-Supervised Learning, SSL)语音模型则提供了涵盖声学与音素层级的分层表示。然而,现有的基于流匹配的语音增强(Speech Enhancement, SE)方法主要在频谱域操作,将SSL特征仅作为外部条件使用,而非直接在SSL潜在空间中建模。为充分挖掘SSL表示的结构丰富性,我们提出PhASE-Flow,一种完全在SSL空间运行的基于流匹配的语音增强框架。该方法建模给定音素表示时清洁声学表示的条件分布,并通过神经声码器重建波形。实验表明,PhASE-Flow在感知质量和可懂度上均优于当前最优基线方法。值得注意的是,其仅需4次采样步骤即可达到竞争性性能,实现高效推理。音频演示可通过 https://anonymous.4open.science/w/phase-flow_demo-E6E1/ 获取。

原文摘要 · Abstract (English)

Flow matching (FM) enables high-fidelity generation, while self-supervised learning (SSL) speech models provide hierarchical representations spanning acoustic and phonetic levels. However, existing FM-based speech enhancement (SE) methods operate primarily in the spectral domain, treating SSL features only as external conditions rather than modeling directly in the SSL latent space. To fully exploit the structural richness of SSL representations, we propose PhASE-Flow, an FM-based SE framework that operates entirely in the SSL space. It models the conditional distribution of clean acoustic representations given phonetic ones, reconstructing the waveform via a neural vocoder. Experiments show that PhASE-Flow outperforms state-of-the-art baselines in perceptual quality and intelligibility. Notably, it achieves competitive performance with only four sampling steps, enabling highly efficient inference. Audio demos are available at https://anonymous.4open.science/w/phase-flow_demo-E6E1/.

语音增强流匹配自监督学习高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。