arXiv:2606.29450eess.AS2026-06中稿 · Interspeech 2026

通过清洁语音引导,提升噪声下语音宽带扩展的清晰度。

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

论文配图:VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion
图 1 · 摘自论文原文
  • 引入速度对比正则与表征对齐,引导生成过程向干净语音靠拢。
  • 在低信噪比下仍保持最低频谱失真和最高语音质量评分。
  • 适合需要高保真语音重建的语音增强与通信场景。

噪声鲁棒的宽带扩展旨在从带噪的低分辨率输入中重建高保真宽带语音。尽管流匹配在语音生成中表现优异,但在噪声干扰下准确恢复干净语音仍具挑战性,主要源于噪声环境下速度估计的不确定性。本文提出 VeRe-Flow,一种基于清洁语音引导的流匹配框架,引入多层次清洁监督以指导生成过程。在速度层面,采用速度对比正则化,使预测速度趋向干净轨迹,同时远离噪声轨迹;在表征层面,引入表征对齐机制,将中间特征与干净语音的自监督学习表征对齐。实验结果表明,该方法在所有基线中达到最低的LSD(平均2.17)和最高的DNSMOS OVRL(6.34),且在生成类方法中取得最高MOS评分(4.58)。

原文摘要 · Abstract (English)

Noise-robust bandwidth expansion aims to reconstruct high-fidelity wideband speech from noisy low-resolution inputs. While flow matching has shown strong performance in speech generation, accurately recovering clean speech from noisy inputs remains challenging due to the ambiguity of velocity estimation under noise. In this work, we propose VeRe-Flow, a clean-guided flow matching framework that introduces multi-level clean supervision to guide the generative process toward clean speech. At the velocity level, we introduce velocity contrastive regularization, which attracts the predicted velocity toward the clean trajectory while repelling it from noisy trajectories. At the representation level, we incorporate representation alignment that aligns intermediate features with clean self-supervised learning representations. The results demonstrate that the proposed method achieves the lowest LSD and highest DNSMOS OVRL among all baselines, and the highest MOS among generative baselines.

语音增强流匹配去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。