Mamba模型在生物声学任务中表现媲美Transformer,显存占用更低。
State Space Models for Bioacoustics: A Comparative Evaluation with Transformers
- 用Mamba架构构建生物声学表示模型BioMamba,采用自监督预训练。
- 在BEANS基准上性能接近顶级Transformer模型AVES,VRAM消耗大幅降低。
- 适合资源受限的野外环境监测场景,计算效率高。
本研究评估了Mamba架构在生物声学中的有效性,提出基于Mamba的音频表征模型BioMamba。我们在大规模音频语料上通过自监督学习对BioMamba进行预训练,并在涵盖多种分类与检测任务的BEANS基准上进行评估。相比当前最先进的Transformer模型AVES,BioMamba在性能上相当,但显著降低了显存占用。结果表明,Mamba可作为真实环境监测中计算高效的替代方案。
原文摘要 · Abstract (English)
In this study, we evaluate the efficacy of the Mamba architecture bioacoustics by introducing BioMamba, a Mamba-based audio representation model for wildlife sounds. We pre-train a BioMamba using self-supervised learning on a large audio corpus and evaluate it on the BEANS benchmark across diverse classification and detection tasks. Compared to the state-of-the-art Transformer-based model (AVES), BioMamba achieves comparable performance while significantly reducing VRAM consumption. Our results demonstrate Mamba's potential as a computationally efficient alternative for real-world environmental monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。