突破16kHz音频限制,用多频段融合提升动物声学识别精度
Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

- 将动物叫声全频谱拆解为多个频段特征,再融合成统一表示
- 在三个数据集上,融合方法比传统基带方法准确率更高
- 适合研究超声波动物行为或需全频段分析的科研人员
动物听觉和发声频率范围远超人类,常延伸至超声波域。然而多数计算生物声学系统依赖16 kHz预训练模型,仅使用0-8 kHz基带,丢弃高频信息。本文研究一种多频段编码框架,将动物叫声全频谱分解为频段特征并融合为统一表示。相似性分析显示,某些编码器生成的频段嵌入具有去相关性,融合后类别分离更优。在三个生物声学数据集上,采用八种预训练模型与五种融合策略的实验表明,融合表示在两个数据集上持续优于基带和时序扩展基线,验证了多频段方法在全频谱动物叫声建模中的潜力。
原文摘要 · Abstract (English)
Animals hear and vocalize across frequency ranges that differ substantially from humans, often extending into the ultrasonic domain. Yet most computational bioacoustics systems rely on audio models pre-trained at 16 kHz, restricting their usable bandwidth to the 0-8 kHz baseband and discarding higher-frequency information present in many bioacoustic recordings. We investigate a multi-band encoding framework that decomposes the full spectrum of animal calls into band features and fuses them into a unified representation. Similarity analyses on models show that certain encoders produce decorrelated band embeddings that improve class separation after fusion. Classification experiments on three bioacoustic datasets using eight pre-trained models and five fusion strategies show that fused representations consistently outperform the baseband and time-expansion baselines on two datasets, showing the potential of multi-band methods for full-spectrum encoding of animal calls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。