arXiv:2511.08261cs.SDcs.LG2025-11中稿 · ICAART 2026

评估鸟类声音分类器的不确定性校准,发现多数模型过自信,小样本校准可显著改善。

Uncertainty Calibration of Multi-Label Bird Sound Classifiers

  • 在BirdSet数据集上系统测试四种多标签分类器的不确定性校准效果
  • 多数模型存在系统性过自信,但罕见类别反而校准更好
  • 仅需少量标注数据,用Platt缩放即可有效提升校准性能

被动声学监测支持大规模生物多样性评估,但可靠的声音分类不仅需要高准确率,还需良好的不确定性估计以支撑决策。生物声学中的校准面临叫声重叠、物种分布长尾及训练与部署数据分布偏移等挑战。本文首次系统评估了四种前沿多标签鸟类声音分类器在BirdSet基准上的校准表现,采用无阈值校准指标(ECE、MCS)和判别性能指标(cmAP),分别评估全局、数据集级和类别级校准。结果显示模型校准性能在不同数据集和类别间差异显著:Perch v2和ConvNeXt$_{BS}$整体校准较好,但普遍存在低估;AudioProtoPNet和BirdMAE则普遍过自信。令人意外的是,罕见类别的校准反而更优。通过简单的后处理校准方法发现,仅需少量标注校准集即可显著改善校准效果,而全局校准参数受数据集差异影响较大。研究强调了在生物声学分类器中评估与改进不确定性校准的重要性。

原文摘要 · Abstract (English)

Passive acoustic monitoring enables large-scale biodiversity assessment, but reliable classification of bioacoustic sounds requires not only high accuracy but also well-calibrated uncertainty estimates to ground decision-making. In bioacoustics, calibration is challenged by overlapping vocalisations, long-tailed species distributions, and distribution shifts between training and deployment data. The calibration of multi-label deep learning classifiers within the domain of bioacoustics has not yet been assessed. We systematically benchmark the calibration of four state-of-the-art multi-label bird sound classifiers on the BirdSet benchmark, evaluating both global, per-dataset and per-class calibration using threshold-free calibration metrics (ECE, MCS) alongside discrimination metrics (cmAP). Model calibration varies significantly across datasets and classes. While Perch v2 and ConvNeXt$_{BS}$ show better global calibration, results vary between datasets. Both models indicate consistent underconfidence, while AudioProtoPNet and BirdMAE are mostly overconfident. Surprisingly, calibration seems to be better for less frequent classes. Using simple post hoc calibration methods we demonstrate a straightforward way to improve calibration. A small labelled calibration set is sufficient to significantly improve calibration with Platt scaling, while global calibration parameters suffer from dataset variability. Our findings highlight the importance of evaluating and improving uncertainty calibration in bioacoustic classifiers.

生物声学不确定性多标签校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。