arXiv:2608.07781eess.AS2026-08

推理时动态修复语音增强中的过抑制问题,提升语音清晰度。

Mitigating Over-Suppression in Speech Enhancement via Inference-Time Rethink-and-Refine Correction Module

论文配图:Mitigating Over-Suppression in Speech Enhancement via Inference-Time Rethink-and-Refine Correction Module
图 1 · 摘自论文原文
  • 用语音识别对齐噪声与增强信号,定位不靠谱的处理段
  • 对可疑段落进行加权混合修复,兼顾听感与语音保真
  • 无需重训练,可适配各类语音增强模型,适合实际部署

我们提出一种推理时重思考与修正模块,解决语音增强模型中常见的过抑制问题——即语音线索与噪声一同被压制。该方法完全在推理阶段运行,无需额外训练,可无缝集成到多种语音增强模型中。通过自动语音识别模型获取单词或音素级对齐,识别出增强不可靠的时间区间,并针对这些区间采用凸插值方式进行选择性重混,每段权重通过优化复合目标函数(平衡感知质量与语音保留)确定。在URGENT 2024和2025、VCTK-DEMAND及MSP-PODCAST数据集上的实验表明,相比传统语音增强方法,该方法在感知质量、语音可懂度和下游任务性能上均取得一致提升,验证了重思考-修正框架在鲁棒语音处理中的有效性。

原文摘要 · Abstract (English)

We present a rethink-and-refine correction module that addresses over-suppression, a common failure mode of speech enhancement (SE) models, where speech cues are suppressed alongside noise. Our method operates entirely in the inference stage without additional training, allowing seamless integration with diverse SE models. Given noisy and enhanced signals, we obtain word- or phoneme-level alignments using an automatic speech recognition model and identify intervals where enhancement is unreliable. These intervals are then selectively remixed through convex interpolation, with per-segment weights optimized to maximize a composite objective balancing perceptual quality and speech preservation. Experiments on the URGENT 2024 and 2025, VCTK-DEMAND, and MSP-PODCAST datasets show consistent improvements in perceptual quality, intelligibility, and downstream performance compared to conventional SE alone, demonstrating the benefit of rethink-and-refine framework for robust speech processing.

语音增强推理优化过抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。