arXiv:2512.10382cs.SD2025-12被引 1

对比三种目标提升流匹配语音增强效率与质量

Investigating training objective for flow matching-based speech enhancement

  • 采用速度预测、x1预测等三种训练目标对比研究
  • 结合感知与信号指标,收敛更快,语音质量显著提升
  • 适合关注语音增强效率与生成质量的工程师

语音增强(SE)旨在从含噪录音中恢复清晰语音。尽管生成方法如得分匹配和薛定谔桥表现出色,但计算成本较高。流匹配通过直接学习将噪声映射到数据的速率场,提供更高效的替代方案。本文系统研究了在三种训练目标下:速度预测、x₁预测和预条件x₁预测,对流匹配在语音增强中的影响,分析其对训练动态与整体性能的作用。此外,引入感知指标(PESQ)和信号指标(SI-SDR),进一步提升了收敛效率与语音质量,在各项评估指标上均取得显著改善。

原文摘要 · Abstract (English)

Speech enhancement(SE) aims to recover clean speech from noisy recordings. Although generative approaches such as score matching and Schrodinger bridge have shown strong effectiveness, they are often computationally expensive. Flow matching offers a more efficient alternative by directly learning a velocity field that maps noise to data. In this work, we present a systematic study of flow matching for SE under three training objectives: velocity prediction, $x_1$ prediction, and preconditioned $x_1$ prediction. We analyze their impact on training dynamics and overall performance. Moreover, by introducing perceptual(PESQ) and signal-based(SI-SDR) objectives, we further enhance convergence efficiency and speech quality, yielding substantial improvements across evaluation metrics.

语音增强流匹配生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。