自适应损失平衡提升鸟类声音多任务分类准确率。
Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

- 设计自适应损失权重机制,解决物种与叫声类型分类不平衡问题。
- 不同预训练模型在不同微调策略下表现各异,最佳方案依赖任务与配置。
- 冻结主干网络微调比全量微调更高效,适合资源受限场景。
被动声学监测中可靠分析鸟类鸣叫需处理多个、不均衡的标注目标。本文扩展BirdCallNet,在长尾分布的WiWa数据集上联合进行物种与叫声类型分类,研究任务损失平衡与预训练表示及适配深度的关系。评估四种鸟类领域编码器(ConvNeXtBS、EAT、BirdMAE、ProtoCLR),在线性探测、注意力探测和全微调三种策略下分别使用独立的物种与叫声头。对比手动固定权重、同方差不确定性加权、动态权重平均(DWA)以及仅在全微调下测试的GradNorm。结果表明,解耦多任务架构对叫声识别提升最稳定,而物种识别效果取决于适配策略。全微调并非始终最优:ConvNeXtBS在线性探测下表现最佳,BirdMAE在注意力探测中胜出。不确定性加权在注意力探测中对物种识别尤为有效,而DWA在全微调中整体更强。GradNorm对部分骨干网络可实现竞争性叫声性能,但物种识别普遍落后且计算与内存开销更高。总体而言,最优损失平衡策略取决于骨干网络、适配方式与目标任务,冻结主干微调比端到端微调更具性能-效率优势。
原文摘要 · Abstract (English)
Reliable analysis of bird vocalisations in passive acoustic monitoring requires models handling multiple, imbalanced annotation targets. We extend BirdCallNet for joint species and call-type classification on the long-tailed WiWa dataset and investigate how task-loss balancing interacts with pretrained representations and adaptation depth. We evaluate four bird-domain encoders, ConvNeXtBS, EAT, BirdMAE, and ProtoCLR, with separate species and call-type heads under linear probing, attentive probing, and full fine-tuning. A manually tuned fixed objective is compared with homoscedastic uncertainty weighting and Dynamic Weight Averaging across all three adaptation regimes, while GradNorm is evaluated only under full fine-tuning. Results indicate that the factorised multi-task formulation yields the most consistent improvements over the combined single-task baseline for call-type recognition, while its effect on species recognition depends on the adaptation regime. Full fine-tuning is not consistently optimal: ConvNeXtBS achieves the highest mean species performance under linear probing, whereas BirdMAE provides the strongest call-type performance under attentive probing. Adaptive weighting benefits species recognition more consistently than call-type recognition. Uncertainty weighting is particularly effective for species recognition under attentive probing, whereas Dynamic Weight Averaging is generally stronger for the same task under full fine-tuning. GradNorm achieves competitive call-type performance for selected backbones but consistently underperforms other weighting strategies for species recognition and incurs higher computational and memory costs. Overall, the preferred loss-balancing strategy depends on the backbone, adaptation regime, and target task, while frozen-backbone adaptation can provide a more favourable performance-efficiency trade-off than end-to-end fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。