轻量级音频模型BioME,让物联网设备也能高效听懂生物声音。
BioME: A Resource-Efficient Bioacoustic Foundational Model for IoT Applications
- 用知识蒸馏压缩大模型,参数减少75%却保持强表现
- 跨域预训练+调制感知特征,提升小设备在复杂环境下的识别力
- 适合低算力物联网场景,开源代码和模型供复用
被动声学监测已成为生物多样性评估、保护与行为生态学的关键策略,尤其在物联网(IoT)设备实现大规模原位音频采集的背景下。尽管近期基于自监督学习(SSL)的音频编码器(如BEATs和AVES)在生物声学任务中表现出色,但其计算开销大且对未见环境鲁棒性不足,限制了在资源受限平台上的部署。本文提出BioME,一种专为生物声学应用设计的资源高效音频编码器。BioME通过层到层的知识蒸馏从高容量教师模型训练,实现强大表征迁移的同时,将参数量减少75%。为进一步提升生态泛化能力,模型在涵盖语音、环境音和动物鸣叫的多领域数据上进行预训练。关键贡献在于引入基于FiLM的调制感知声学特征,注入受信号处理启发的先验知识,增强低容量情况下的特征解耦能力。在多个生物声学任务中,BioME性能达到或超越更大模型(包括其教师模型),同时适用于资源受限的IoT部署。为保障可复现性,代码与预训练检查点已公开。
原文摘要 · Abstract (English)
Passive acoustic monitoring has become a key strategy in biodiversity assessment, conservation, and behavioral ecology, especially as Internet-of-Things (IoT) devices enable continuous in situ audio collection at scale. While recent self-supervised learning (SSL)-based audio encoders, such as BEATs and AVES, have shown strong performance in bioacoustic tasks, their computational cost and limited robustness to unseen environments hinder deployment on resource-constrained platforms. In this work, we introduce BioME, a resource-efficient audio encoder designed for bioacoustic applications. BioME is trained via layer-to-layer distillation from a high-capacity teacher model, enabling strong representational transfer while reducing the parameter count by 75%. To further improve ecological generalization, the model is pretrained on multi-domain data spanning speech, environmental sounds, and animal vocalizations. A key contribution is the integration of modulation-aware acoustic features via FiLM conditioning, injecting a DSP-inspired inductive bias that enhances feature disentanglement in low-capacity regimes. Across multiple bioacoustic tasks, BioME matches or surpasses the performance of larger models, including its teacher, while being suitable for resource-constrained IoT deployments. For reproducibility, code and pretrained checkpoints are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。