arXiv:2508.10436cs.SDcs.AI2025-08

用交替处理法减少语音增强中的失真,提升音质。

Alternating Approach-Putt Models for Multi-Stage Speech Enhancement

  • 设计后处理网络PuttNet,与主增强模型交替使用
  • 在PESQ、STOI和CBAK指标上显著提升语音质量
  • 适合需要高保真语音输出的场景,如语音识别前处理

基于人工神经网络的语音增强旨在去除噪声同时保留语音内容。然而,现有语音增强网络常引入失真,称为伪影,降低音频质量。本文提出一种后处理神经网络,用于缓解语音增强模型引入的伪影。受高尔夫中‘攻球’后接‘推杆’的启发,将该模型命名为PuttNet。实验表明,交替使用语音增强模型与提出的Putt模型,可显著提升语音感知质量(PESQ)、客观可懂度(STOI)及背景噪声侵入度(CBAK)评分。此外,通过图示分析说明,交替策略优于单一模型重复应用。

原文摘要 · Abstract (English)

Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as artifacts, which can degrade audio quality. In this work, we propose a post-processing neural network designed to mitigate artifacts introduced by speech enhancement models. Inspired by the analogy of making a `Putt' after an `Approach' in golf, we name our model PuttNet. We demonstrate that alternating between a speech enhancement model and the proposed Putt model leads to improved speech quality, as measured by perceptual quality scores (PESQ), objective intelligibility (STOI), and background noise intrusiveness (CBAK) scores. Furthermore, we illustrate with graphical analysis why this alternating Approach outperforms repeated application of either model alone.

语音增强降噪后处理伪影消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。