无干净数据下实现动物叫声去噪,用伪干净样本训练模型。
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
- 用语音增强模型预处理生成伪干净叫声作为训练目标。
- 在跨物种噪声数据集上表现接近现有方法,验证了有效性。
- 适合生态监测、动物行为研究等需要声音分析的领域。
动物叫声去噪任务与人类语音增强类似,但面临更复杂的发声机制和录音环境多样性挑战。由于缺乏大规模多样化的干净叫声数据集,现有方法受限。本文提出使用伪干净目标(即通过语音增强模型预去噪的叫声)和不含叫声的背景噪声片段作为训练数据。构建了一个涵盖多种物种、声学环境和地理区域的生物声学训练集,并引入一个非重叠基准测试集,包含不同类群的干净叫声和噪声样本。实验表明,基于demucs和CleanUNet的去噪模型在伪干净数据上训练后,在基准测试集上取得具有竞争力的结果。相关数据、代码、库和演示已公开于https://earthspecies.github.io/biodenoising/。
原文摘要 · Abstract (English)
Animal vocalization denoising is a task similar to human speech enhancement, which is relatively well-studied. In contrast to the latter, it comprises a higher diversity of sound production mechanisms and recording environments, and this higher diversity is a challenge for existing models. Adding to the challenge and in contrast to speech, we lack large and diverse datasets comprising clean vocalizations. As a solution we use as training data pseudo-clean targets, i.e. pre-denoised vocalizations, and segments of background noise without a vocalization. We propose a train set derived from bioacoustics datasets and repositories representing diverse species, acoustic environments, geographic regions. Additionally, we introduce a non-overlapping benchmark set comprising clean vocalizations from different taxa and noise samples. We show that that denoising models (demucs, CleanUNet) trained on pseudo-clean targets obtained with speech enhancement models achieve competitive results on the benchmarking set. We publish data, code, libraries, and demos at https://earthspecies.github.io/biodenoising/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。