arXiv:2503.14854eess.AS2025-03被引 2

用噪声信号替代干净信号,实现无需干净数据的语音增强。

Analysis and Extension of Noisy-target Training for Unsupervised Target Signal Enhancement

  • 以噪声目标信号代替干净信号进行训练,简化无监督语音增强流程。
  • 在少量干净数据下仍能有效提升语音质量,证明方法鲁棒性。
  • 扩展至去混响和去削波任务,适用范围更广,适合资源受限场景。

基于深度神经网络的目标信号增强(TSE)通常依赖于干净目标信号进行有监督训练,但获取干净信号成本高且不总是可用。因此,发展不依赖干净信号的无监督方法至关重要。在众多无监督TSE方法中,噪声目标训练(NyTT)已成为基础方法:简单地将监督训练中的干净目标替换为噪声目标,实验已证明其可实现有效增强。尽管效果显著且操作简单,其内在机制与具体行为仍不明确。本文从多角度分析NyTT,实验揭示其作用机制、理想条件及在少量干净信号下的有效性。进一步,基于分析结果提出改进版NyTT,并验证其在去混响与去削波任务中的能力,拓展了该方法的应用边界。

原文摘要 · Abstract (English)

Deep neural network-based target signal enhancement (TSE) is usually trained in a supervised manner using clean target signals. However, collecting clean target signals is costly and such signals are not always available. Thus, it is desirable to develop an unsupervised method that does not rely on clean target signals. Among various studies on unsupervised TSE methods, Noisy-target Training (NyTT) has been established as a fundamental method. NyTT simply replaces clean target signals with noisy ones in the typical supervised training, and it has been experimentally shown to achieve TSE. Despite its effectiveness and simplicity, its mechanism and detailed behavior are still unclear. In this paper, to advance NyTT and, thus, unsupervised methods as a whole, we analyze NyTT from various perspectives. We experimentally demonstrate the mechanism of NyTT, the desirable conditions, and the effectiveness of utilizing noisy signals in situations where a small number of clean target signals are available. Furthermore, we propose an improved version of NyTT based on its properties and explore its capabilities in the dereverberation and declipping tasks, beyond the denoising task.

语音增强无监督学习降噪去混响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。