用声纹轮廓合成训练数据,提升小鼠叫声去噪效果
Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

- 基于声调轮廓生成带噪训练数据,构建时频域复数比值掩码模型
- 在真实野外录音中显著改善基频与泛音轮廓追踪能力
- 适合需要高精度生物叫声分析的神经科学与行为学研究者
生物声学记录常受环境噪声干扰,导致微弱或被噪声掩盖的发声难以分析。尽管卷积神经网络(尤其是U-Net)在语音和音乐去噪中表现优异,但其在生物声信号上的应用受限于清洁数据稀缺。为此,本文提出一种训练集合成方法,开发了一种监督去噪模型,该模型在时频域预测复数比值掩码。模型利用表示基频及一个或多个谐波成分的声纹轮廓(ridges),既用于训练数据合成,也用于设计加权损失函数(ridge-guided loss),使网络在去噪时更关注声纹区域,从而更好保留发声细节。以家鼠超声发声(USVs)为例,该方法在真实野外录音中相比以往信号处理方法显著提升了基频与谐波轮廓追踪能力。此外,基于去噪数据训练的分类器在未见的野生与家养鼠噪声录音上,分类性能优于基于原始噪声数据训练的模型。在合成测试数据上,该方法在广泛输入信噪比条件下均显著提升尺度不变信号失真比。尽管聚焦于超声发声,该方法可推广至其他具有可追踪声纹轮廓的生物声信号,实现基于声纹的训练集合成与去噪。
原文摘要 · Abstract (English)
Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。