arXiv:2605.14031cs.SDcs.CV2026-05被引 1

在有限标注的生物声学数据上,大尺度预训练比精细设计更有效。

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

论文配图:Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
图 1 · 摘自论文原文
  • 用掩码自编码器在多样音频上预训练,再迁移到生物声学分类任务。
  • 在小规模数据下,领域特定数据进一步预训练反而降低性能。
  • 适合资源有限但需高精度物种识别的研究者参考。

生物声学识别需细粒度声学理解以区分相似声音的物种。然而,iNaturalist等大型数据集常仅提供每段录音一个正类标签,导致监督学习困难。受计算机视觉启发,近年研究转向自监督学习以捕捉音频内在结构,无需详尽标注。掩码自编码器(MAE)在大规模音频数据上表现优异,但在中等规模生物声学场景下的效果仍不明确。本文系统研究了在iNatSounds数据集上使用MAE进行物种分类的预训练策略,分析了预训练数据规模、领域专属性、数据清洗和迁移策略的影响。结果表明:尽管通用音频数据预训练表现最佳,但额外在领域特定数据上进行掩码重建预训练收益甚微,甚至可能劣于现成模型;在数据总量有限时,选择性数据过滤也几乎无益。研究显示,在中等规模细粒度生物声学任务中,预训练规模远比目标设计重要。这些发现明确了MAE预训练的有效边界,为弱监督条件下的模型选择提供了实用指导。

原文摘要 · Abstract (English)

Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositories such as iNaturalist are weakly annotated, often with only a single positive species label per recording, making supervised learning particularly challenging. Inspired by advances in computer vision, recent approaches have shifted toward self-supervised learning to capture the underlying structure of audio without relying on exhaustive annotations. In particular, masked autoencoders (MAE) have shown strong transferability on massive audio corpora, yet their effectiveness in more modest bioacoustic settings remains underexplored. In this work, we conduct a systematic study of MAE pretraining for species classification on iNatSounds, analyzing the impacts of pretraining data scale, domain specificity, data curation, and transfer strategies. Consistent with prior work, we find that models pretrained on diverse general audio data achieve the best transfer performance on iNatSounds. Contrary to observations from large-scale audio benchmarks, we find that (1) additional masked reconstruction pretraining on domain-specific data provides limited benefits and may even degrade performance relative to off-the-shelf models, and (2) selective data filtering offers a negligible advantage when the overall data scale is limited. Our results indicate that, in moderate-sized fine-grained bioacoustic settings, pretraining scale dominates objective design. These findings further clarify when MAE-based pretraining is effective and provide practical guidance for model selection under limited supervision.

生物声学自监督学习掩码自编码器小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。