解析生物声学嵌入中的语音特征,提升模型透明度与适用性
Beyond task performance: Decoding bioacoustic embeddings with speech features

- 用线性与非线性探测器分析模型编码的语音特征
- 响度特征恢复效果最好(R²=0.76),基频最难恢复(R²=0.33)
- 提出基于特征可恢复性的模型选择指南,适合稀有物种研究
预训练音频嵌入在生物声学中已成标准,但对其编码的声学特征及任务相关性了解甚少,阻碍了模型透明度与在稀有物种或数据稀缺场景下的应用。本文通过在六个分类群中使用88维eGeMAPS特征,采用线性与非线性回归探测器,量化各模型捕捉的声学属性。结果表明不存在‘免费午餐’:无单一模型能覆盖全部特征空间。拼接嵌入表现最佳,说明不同模型覆盖互补的声学空间。响度特征恢复能力最强(R²=0.76),而基频最差(R²=0.33)。结合每物种特征显著性(NMI)与可恢复性,提出面向生物声学的数据驱动模型选择策略。
原文摘要 · Abstract (English)
Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species or data-scarce domains. Here we reveal which speech-like features are encoded in bioacoustic representations. Using the 88~eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a ``no free lunch'' pattern: no single model captures the full feature space. A concatenated embedding achieves the highest performance, suggesting complementary acoustic space coverage across models. Loudness features are best encoded ($R^2 = 0.76$) while F0 is hardest to recover ($R^2 = 0.33$). By cross-referencing recoverability with per-species feature salience (NMI), we derive data-driven model selection guidance for bioacoustics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。