通过校准增强提升跨方言鸟鸣识别准确率,兼顾模型透明性与实用性。
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
- 用频率敏感归一化和对抗训练学习区域无关声学特征
- 跨方言识别准确率最高提升20个百分点,且本地性能不下降
- 适合生态监测场景,对模型解释性要求高的研究者
被动声学监测采集的鸟鸣存在方言变异问题,影响自动识别。本文在包含三个区域、十种鸟类的8秒音频数据集DB3V上提出一种可部署框架,基于时间延迟神经网络(TDNN)。通过实例频率归一化与门控松弛型频率归一化,结合梯度反转对抗训练,实现区域无关嵌入学习。采用多层级增强策略:波形扰动、稀有类别Mixup,以及基于CycleGAN的区域2(内陆平原)风格音频合成,并引入方言校准增强(DCA),软性降低合成样本权重以减少伪影。整体系统在跨方言识别上较基线TDNN最高提升20个百分点,同时保持区域内性能。Grad-CAM与LIME分析显示,鲁棒模型聚焦于稳定的谐波频带,提供生态意义明确的解释。结果表明,轻量、透明且抗方言干扰的鸟声识别是可行的。
原文摘要 · Abstract (English)
Dialect variation hampers automatic recognition of bird calls collected by passive acoustic monitoring. We address the problem on DB3V, a three-region, ten-species corpus of 8-s clips, and propose a deployable framework built on Time-Delay Neural Networks (TDNNs). Frequency-sensitive normalisation (Instance Frequency Normalisation and a gated Relaxed-IFN) is paired with gradient-reversal adversarial training to learn region-invariant embeddings. A multi-level augmentation scheme combines waveform perturbations, Mixup for rare classes, and CycleGAN transfer that synthesises Region 2 (Interior Plains)-style audio, , with Dialect-Calibrated Augmentation (DCA) softly down-weighting synthetic samples to limit artifacts. The complete system lifts cross-dialect accuracy by up to twenty percentage points over baseline TDNNs while preserving in-region performance. Grad-CAM and LIME analyses show that robust models concentrate on stable harmonic bands, providing ecologically meaningful explanations. The study demonstrates that lightweight, transparent, and dialect-resilient bird-sound recognition is attainable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。