arXiv:2602.09295cs.LGcs.SD2026-02

用主动学习方法从30年声学数据中挖掘出最大规模的虎鲸音频数据集。

Positive-Unlabelled Active Learning to Curate a Dataset for Orca Resident Interpretation

  • 采用弱监督正负样本主动学习策略,自动识别海洋哺乳动物叫声
  • 构建的Transformer模型在多个数据集上准确率超现有水平,能效更优
  • 数据集覆盖超900小时虎鲸音频,适合濒危种群保护与生态研究

本研究首次系统整理了超过30年、覆盖南部分布区虎鲸(SRKW)栖息地的全部公开水下麦克风数据,构建了迄今为止最大的虎鲸声学数据集。通过弱监督正负样本主动学习策略,自动识别海洋哺乳动物发声事件。基于Transformer的有无存在分类器在4个专家标注数据集中,3个达到最优性能,且能耗更低。多类物种分类器在DCLDE-2026数据集上达53.2%的top-1准确率(11个训练类别,4个测试类别),生态型分类器达33.6%(4个训练类别,5个测试类别)。最终产出包括919小时虎鲸数据、230小时Bigg's虎鲸数据、1374小时未标注生态型虎鲸数据、1501小时座头鲸数据、88小时海狮数据、246小时太平洋白侧海豚数据,以及超过784小时未明确物种的数据。该数据集规模超越DCLDE-2026、Ocean Networks Canada和OrcaSound之和。标签以CC-BY 4.0开源,音频数据遵循原始来源许可。其全面性适用于无监督机器翻译、栖息地使用分析及濒危生态型保护工作。

原文摘要 · Abstract (English)

This work presents the largest curation of Southern Resident Killer Whale (SRKW) acoustic data to date, also containing other marine mammals in their environment. We systematically search all available public archival hydrophone data within the SRKW habitat (over 30 years of audio data). The search consists of a weakly-supervised, positive-unlabelled, active learning strategy to identify all instances of marine mammals. The resulting transformer-based presence or absence classifiers outperform state-of-the-art classifiers on 3 of 4 expert-annotated datasets in terms of accuracy and energy efficiency. The fleet of WHISPER detection models range from 0.58 (0.48-0.67) AUROC with WHISPER-tiny to 0.77 (0.63-0.93) with WHISPER-large-v3. Our multiclass species classifier obtains a top-1 accuracy of 53.2\% (11 train classes, 4 test classes) and our ecotype classifier obtains a top-1 accuracy of 33.6\% (4 train classes, 5 test classes) on the DCLDE-2026 dataset. We yield 919 hours of SRKW data, 230 hours of Bigg's orca data, 1374 hours of orca data from unlabelled ecotypes, 1501 hours of humpback data, 88 hours of sea lion data, 246 hours of pacific white-sided dolphin data, and over 784 hours of unspecified marine mammal data. This SRKW dataset is larger than DCLDE-2026, Ocean Networks Canada, and OrcaSound combined. The curated species labels are available under CC-BY 4.0 license, and the corresponding audio data are available under the licenses of the original owners. The comprehensive nature of this dataset makes it suitable for unsupervised machine translation, habitat usage surveys, and conservation endeavours for this critically endangered ecotype.

声学监测虎鲸保护主动学习生态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。