arXiv:2607.03221eess.AScs.SD2026-07中稿 · IWAENC 2026

用分离音频提升鸟类识别,减少误检。

Mixture-Constrained Max Pooling Improves Separation-Based Bird Species Classification

论文配图:Mixture-Constrained Max Pooling Improves Separation-Based Bird Species Classification
图 1 · 摘自论文原文
  • 用两种分离模型集成,先分音再分类。
  • 新方法使有鸟时检出率升12.3%,无鸟时误报降18.7%。
  • 适合做野外鸟类监测的团队或研究者。

从野外录音中进行鸟类物种分类仍具挑战性,主要源于声音重叠和不完整的物种标签。本文将语音分离作为鸟类分类的预处理步骤,以提升多物种检测性能。具体地,采用两个分离器(FTRNN 和 TF-Locoformer)的集成,二者均通过混合不变训练(MixIT)进行训练。为缓解分离错误导致的误报增加问题,提出混合约束最大池化(MCM),该方法根据原始混合信号中各物种的概率,对每个分离通道的预测概率进行裁剪。分类器分别应用于每个分离输出及原始混合信号,并通过 MCM 将预测结果聚合为最终的每物种概率。在两个真实数据集上的实验表明,集成分离器优于单一分离器,且 MCM 在多个指标上优于标准最大池化。结果揭示,分离不仅能提升实际存在物种的真阳性率,还能降低不存在物种的假阳性率。

原文摘要 · Abstract (English)

Bird species classification from field recordings remains challenging due to overlapping vocalizations and incomplete species labels. We study source separation as a preprocessing for bird species classification to improve multi-species detection. Specifically, we employ an ensemble of two separators, FTRNN and TF-Locoformer, both trained with mixture invariant training (MixIT). To address the false positive gain caused by separation errors in separated outputs, we propose mixture-constrained max pooling (MCM), which clips the predicted probability from each separated channel based on the corresponding species probability in the original mixture. The classifier is applied to each separated output and the original mixture independently, and MCM aggregates the predictions into a final per-species probability. Experiments on two real-world datasets show that the ensemble outperforms individual separators and MCM outperforms standard max pooling across multiple metrics, and reveal that separation leads to both true positive gain for present species and false positive gain for absent species.

鸟类识别语音分离多物种检测深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。