arXiv:2512.03563cs.SDcs.AI2025-12

Mamba模型在生物声学任务中表现媲美Transformer,显存占用更低。

State Space Models for Bioacoustics: A Comparative Evaluation with Transformers

  • 用Mamba架构构建生物声学表示模型BioMamba,采用自监督预训练。
  • 在BEANS基准上性能接近顶级Transformer模型AVES,VRAM消耗大幅降低。
  • 适合资源受限的野外环境监测场景,计算效率高。

本研究评估了Mamba架构在生物声学中的有效性,提出基于Mamba的音频表征模型BioMamba。我们在大规模音频语料上通过自监督学习对BioMamba进行预训练,并在涵盖多种分类与检测任务的BEANS基准上进行评估。相比当前最先进的Transformer模型AVES,BioMamba在性能上相当,但显著降低了显存占用。结果表明,Mamba可作为真实环境监测中计算高效的替代方案。

原文摘要 · Abstract (English)

In this study, we evaluate the efficacy of the Mamba architecture bioacoustics by introducing BioMamba, a Mamba-based audio representation model for wildlife sounds. We pre-train a BioMamba using self-supervised learning on a large audio corpus and evaluate it on the BEANS benchmark across diverse classification and detection tasks. Compared to the state-of-the-art Transformer-based model (AVES), BioMamba achieves comparable performance while significantly reducing VRAM consumption. Our results demonstrate Mamba's potential as a computationally efficient alternative for real-world environmental monitoring.

Mamba生物声学自监督高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。