arXiv:2503.00296cs.SDcs.LG2025-03被引 5

用合成数据训练模型,实现零样本生物声学事件检测。

Synthetic data enables context-aware bioacoustic sound event detection

  • 用合成数据构建多样化声景,带强时序标签。
  • 8.8千小时数据训练,少样本检测性能提升64%。
  • 公开13个任务基准,适合生态学家直接使用。

我们提出一种训练基础模型的方法,增强其在生物声学信号处理领域的上下文学习能力。通过合成生成的训练数据,采用基于领域随机化的流水线,构建具有强时序标签的多样化声景。生成超过8.8千小时的强标签音频,并训练一个基于查询示例的Transformer模型,实现少样本生物声学事件检测。第二项贡献是公开13个多样化的少样本生物声学任务基准。模型表现优于已有方法,相比其他无需训练的方法相对提升64%。我们证明这得益于模型规模、数据量以及算法改进。我们通过API提供训练好的模型,为生态学家和行为学家提供免训练的生物声学事件检测工具。

原文摘要 · Abstract (English)

We propose a methodology for training foundation models that enhances their in-context learning capabilities within the domain of bioacoustic signal processing. We use synthetically generated training data, introducing a domain-randomization-based pipeline that constructs diverse acoustic scenes with temporally strong labels. We generate over 8.8 thousand hours of strongly-labeled audio and train a query-by-example, transformer-based model to perform few-shot bioacoustic sound event detection. Our second contribution is a public benchmark of 13 diverse few-shot bioacoustics tasks. Our model outperforms previously published methods, and improves relative to other training-free methods by $64\%$. We demonstrate that this is due to increase in model size and data scale, as well as algorithmic improvements. We make our trained model available via an API, to provide ecologists and ethologists with a training-free tool for bioacoustic sound event detection.

生物声学合成数据少样本检测零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。