arXiv:2607.22458cs.LGcs.AI2026-07

音频大模型能自动捕捉物种演化关系,无需专门训练

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

论文配图:Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining
图 1 · 摘自论文原文
  • 用语音嵌入空间中的距离推断物种演化树
  • 鲸类语音与演化树相关性高达0.82,远超手工特征
  • 专用于鸟类的模型反而不如通用模型,说明领域预训练非必需

我们测试了四种大型预训练音频模型(AST、CLAP、BEATs-bio 和 BirdNET)在未见任务上的表现:从物种发声中恢复系统发育距离。若嵌入空间几何结构反映生命之树,则表示学习到了超越标签的信息。在32种海洋哺乳动物(1,754段录音,来自Watkins海洋哺乳动物声音数据库)中,26种鲸类的模型表现优异(CLAP r=0.82,BEATs-bio r=0.82,AST r=0.74;均p<0.001),相关性为已报道最高之一。手工提取的MFCC特征(105维)无显著相关性(r=0.040,p=0.338)。该差距在主成分降至105维后仍存在,排除维度影响。控制主导频率后部分Mantel检验显示(partial Mantel r=0.404,保留97%方差),结果并非仅由音高导致。对20种鸟类(使用Jetz等,2012年系统发育树)重复实验,加入在约6,000种鸟类上端到端训练的BirdNET模型,通用模型依然表现更优(AST r=0.55,CLAP r=0.52)。意外的是,尽管生物声学专用的BEATs-bio和针对鸟类的BirdNET表现均较差(相关性约0.32–0.36),说明匹配训练域并不能带来优势。预训练音频嵌入能在两个独立辐射谱系中捕捉演化信息,而领域特定预训练并非必要条件。

原文摘要 · Abstract (English)

Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance from species vocalizations. If the geometry of the embedding space tracks the tree of life, the representation is picking up something deeper than the labels the model was optimized for. We run Mantel tests across two independent radiations. In 32 marine mammal species (1,754 recordings from the Watkins Marine Mammal Sound Database) the foundation models recover strong phylogenetic signal within the 26 cetaceans (CLAP r=0.82, BEATs-bio r=0.82, AST r=0.74; all p<0.001), among the highest acoustic-phylogenetic correlations reported for any taxon. Hand-crafted MFCC features (105d) find nothing (r=0.040, p=0.338). The gap survives after PCA-projecting every embedding down to 105 dimensions, so it is not an artefact of representation size. It also survives a partial Mantel test controlling for dominant frequency (partial Mantel r=0.404, keeping 97% of the variance explained), so it is not just pitch in disguise. We repeat the analysis on 20 bird species using the Jetz et al. (2012) phylogeny, and this time add BirdNET, a classifier trained end-to-end on around 6,000 bird species. The general-purpose foundation models recover the signal again (AST r=0.55, CLAP r=0.52). The unexpected result is that neither BirdNET nor the bioacoustic BEATs-bio beat them (r around 0.32 to 0.36). Matching the training domain to the target taxon does not, by itself, help. Pretrained audio embeddings carry evolutionary information across two independent radiations, and domain-specific pretraining is not required for it to emerge.

音频模型演化信号自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。