arXiv:2509.00405cs.SDeess.AS2025-09中稿 · InterSpeech2025

提出场景感知判别器,提升语音增强在不同环境下的表现

SaD: A Scenario-Aware Discriminator for Speech Enhancement

  • 设计场景感知判别器,捕捉不同场景的特征并进行频域划分
  • 在两个数据集上验证,显著提升多种生成模型的语音增强效果
  • 无需改动生成器结构,适配性强,适合多场景语音处理应用

基于生成对抗网络的模型在语音增强领域表现出色。然而,当前优化策略主要聚焦于生成器架构改进或判别器质量评估指标提升,忽视了不同场景中蕴含的丰富上下文信息。本文提出一种场景感知判别器,通过捕获场景特异性特征并进行频域划分,实现对生成语音质量更准确的评估。我们在三个代表性模型上使用两个公开数据集进行了全面实验。结果表明,该方法无需改变生成器结构即可有效适应多种架构,在不同场景下均能进一步提升语音增强性能。

原文摘要 · Abstract (English)

Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models predominantly focus on refining the architecture of the generator or enhancing the quality evaluation metrics of the discriminator. This approach often overlooks the rich contextual information inherent in diverse scenarios. In this paper, we propose a scenario-aware discriminator that captures scene-specific features and performs frequency-domain division, thereby enabling a more accurate quality assessment of the enhanced speech generated by the generator. We conducted comprehensive experiments on three representative models using two publicly available datasets. The results demonstrate that our method can effectively adapt to various generator architectures without altering their structure, thereby unlocking further performance gains in speech enhancement across different scenarios.

语音增强GAN场景感知判别器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。