用不确定性加权融合声音与频谱特征,提升动物叫声分类准确率。
Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior

- 双流架构分别处理波形和频谱,动态计算各模态置信度并加权融合。
- 在猪和狗的叫声数据集上,准确率分别达59.4%和73.1%,宏F1提升超20%。
- 无需标注可靠性,对噪声、混响等干扰鲁棒,适合野外监测场景。
人工智能在动物声学监测中潜力巨大,可用于畜牧精准管理、野生动物保护与生态研究,因叫声能更早且低成本地反映健康、压力与社会状态。但实际录音常受环境噪声、混响、重叠叫声及传感器退化影响,导致自动分类困难。现有两种声学表征各有局限:原始波形保留时间微结构但易受截断与混响破坏,而log-Mel频谱图捕捉谐波结构却丢失相位信息且对宽带噪声敏感。为此,本文提出不确定性感知融合(UAF),一种双流框架,为每种表征估计高斯不确定性,并通过不确定性加权融合。该机制自动赋予更可信模态更高权重,无需可靠性标签。在跨物种、基于个体身份的评估中(训练时未见个体),采用均值池化的UAF在17类SoundWel猪叫声基准上取得59.4%准确率与39.7%宏F1;在3类DogBark数据集上达73.1%准确率与71.5%宏F1,优于静态拼接融合15.7%与20.4%相对宏F1。消融实验显示,不确定性融合是性能提升的主要原因,而非语音的时间特性。
原文摘要 · Abstract (English)
Artificial intelligence offers substantial potential for acoustic monitoring of animals, from welfare assessment in precision livestock farming to wildlife conservation and ecological research, where vocalizations can indicate health, stress, and social states earlier and at lower cost than manual observation. However, recordings in these settings are obtained under uncontrolled conditions, including environmental noise, reverberation, overlapping calls, and sensors that degrade without notice. As a consequence, automated classification of animal vocalizations remains challenging, and the two dominant acoustic representations show complementary limitations: raw waveforms preserve temporal microstructure but degrade under clipping and reverberation, while log-Mel spectrograms capture harmonic organization but lose phase information and are sensitive to broadband noise. To address these challenges, we propose Uncertainty-Aware Fusion (UAF), a dual-stream framework that estimates Gaussian uncertainty for each representation and fuses them via uncertainty weighting. This mechanism assigns greater weight to the more confident representation with no reliability labels required. In a cross-species, identity-based evaluation excluding all individuals seen during training, UAF (mean pooling) achieves 59.4\% accuracy / 39.7\% macro F1 on the 17-class SoundWel pig vocalization benchmark and 73.1\% accuracy / 71.5\% macro F1 on the 3-class DogBark dataset, outperforming static-concatenation fusion by 15.7\% and 20.4\% relative macro F1, respectively. Ablations over four temporal aggregation strategies show that uncertainty fusion, rather than the temporal characteristics of animal calls, is the primary driver of the performance gain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。