arXiv:2507.07631eess.AScs.SD2025-07中稿 · Frontiers in signa…被引 3

用自监督特征空间距离优化语音增强,提升多任务泛化能力

Generic Speech Enhancement with Self-Supervised Representation Space Loss

  • 在自监督模型的特征空间中最小化增强语音与干净语音的距离
  • 在多个下游任务上性能提升,同时保持语音听感质量
  • 适合需要通用语音增强模块的多场景应用

单通道语音增强广泛用于缓解干扰信号影响。传统方法需为每项任务单独调优,导致模型难以泛化到未知下游任务。本文旨在构建一个通用语音增强前端,以提升多种下游任务的表现。为此,提出一种新训练准则:在自监督学习模型的特征表示空间中,最小化增强语音与真实干净语音之间的距离。由于自监督特征能有效表达对多种下游任务有用的高层语音信息,该方法有望使语音增强模型保留此类信息。实验验证表明,该方法在保持语音感知质量的同时,显著提升了多个语音任务的性能。

原文摘要 · Abstract (English)

Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.

语音增强自监督泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。