用噪声录音生成高保真鸟鸣,还能精准控制物种和描述。
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
- 先增强信号再扩散生成,分两步提升音质。
- 噪声抑制效果超现有方法,音质指标大幅提升。
- 适合生态监测与濒危物种数据补充场景。
生成模型为生物声学带来新机遇,可合成逼真动物鸣叫,助力生物监测并补充濒危物种稀缺数据。然而,直接从嘈杂野外录音生成鸟鸣波形仍是重大挑战。我们提出BirdDiff框架,基于12种野生鸟类的噪声数据集,合成鸟鸣。模型包含“零阶层”多尺度自适应增强阶段,随后是基于三模态条件(梅尔倒谱系数、物种标签、文本描述)的扩散生成器。增强阶段在提升信噪比(SNR)的同时最小化频谱失真,相比三种常用无训练增强方法,实现最高SNR提升(+10.45 dB)和最低Itakura-Saito距离(0.54)。对比基线生成模型DiffWave,BirdDiff在生成质量上显著提升:弗雷切特音频距离(0.590 → 0.213)、Jensen-Shannon散度(0.259 → 0.226)、统计差异频段数(7.33 → 5.58)。为评估物种细节保留能力,使用在原始数据集上训练的ResNet50分类器识别生成样本,准确率从35.9%(DiffWave)提升至70.1%(BirdDiff),其中8种物种超过70%准确率。结果表明,BirdDiff能直接从噪声录音生成高保真、可控的鸟鸣。
原文摘要 · Abstract (English)
Generative modeling offers new opportunities for bioacoustics, enabling the synthesis of realistic animal vocalizations that could support biomonitoring efforts and supplement scarce data for endangered species. However, directly generating bird call waveforms from noisy field recordings remains a major challenge. We propose BirdDiff, a generative framework designed to synthesize bird calls from a noisy dataset of 12 wild bird species. The model incorporates a "zeroth layer" stage for multi-scale adaptive bird-call enhancement, followed by a diffusion-based generator conditioned on three modalities: Mel-frequency cepstral coefficients, species labels, and textual descriptions. The enhancement stage improves signal-to-noise ratio (SNR) while minimizing spectral distortion, achieving the highest SNR gain (+10.45 dB) and lowest Itakura-Saito Distance (0.54) compared to three widely used non-training enhancement methods. We evaluate BirdDiff against a baseline generative model, DiffWave. Our method yields substantial improvements in generative quality metrics: Fréchet Audio Distance (0.590 to 0.213), Jensen-Shannon Divergence (0.259 to 0.226), and Number of Statistically-Different Bins (7.33 to 5.58). To assess species-specific detail preservation, we use a ResNet50 classifier trained on the original dataset to identify generated samples. Classification accuracy improves from 35.9% (DiffWave) to 70.1% (BirdDiff), with 8 of 12 species exceeding 70% accuracy. These results demonstrate that BirdDiff enables high-fidelity, controllable bird call generation directly from noisy field recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。