无需训练的音频内容识别基准,评估模型在零样本下的语义对齐能力
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
- 用冻结的音频特征+无标签PCA校正,测试嵌入空间的几何对齐性
- 在12.5万段音频上验证,跨领域表现稳定,但低资源语言检索失效
- 适合评估模型零样本泛化能力,尤其关注语音与生物声学应用
通用音频表征旨在将同一事件的不同声学实例映射到相近点,实现零样本内容识别。不同于依赖参数更新的监督分类基准,我们提出VocSim——一个无需训练、不使用标签的基准,仅通过每子集的无标签PCA白化纠正各向异性。VocSim整合了来自19个数据集的12.5万段单源音频,涵盖人声、动物鸣叫和环境音,排除混响干扰。采用Precision@k衡量局部纯度,用全局分离率(GSR)评估点级分类分离性,并以置换基线校准性能提升。仅用冻结的Whisper特征、时频池化与无标签PCA即可取得强零样本表现,跨领域GSR排名稳定(Kendall's tau=0.60)。但在盲测低资源语言(Shipibo-Conibo、Chintang)中,局部检索崩溃但仍优于随机,暴露出跨语言泛化差距。外部验证显示,最优嵌入可预测鸟类感知相似性,提升生物声学分类,且在HEAR基准上达领先水平。数据、代码与公开排行榜已发布。
原文摘要 · Abstract (English)
General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability via parameter updates, we introduce VocSim, a training-free benchmark probing the intrinsic geometric alignment of frozen embeddings, with no parameters updated and no labels used (a label-free PCA whitening is fit per subset to correct anisotropy). VocSim aggregates 125k single-source clips from 19 corpora spanning human speech, animal vocalizations, and environmental sounds, isolating content representation from source separation (polyphonic mixtures are out of scope). We evaluate embeddings with Precision@k for local purity and the Global Separation Rate (GSR) for point-wise class separation, calibrated by lift over an empirical permutation baseline. A simple pipeline of frozen Whisper features, time-frequency pooling, and label-free PCA yields strong zero-shot performance with stable GSR rankings across domains (Kendall's tau = 0.60). However, on blind low-resource speech (Shipibo-Conibo, Chintang), local retrieval collapses while remaining above chance, exposing a cross-lingual speech generalization gap. As external validation, our top embeddings predict avian perceptual similarity, improve bioacoustic classification, and achieve state-of-the-art on the HEAR benchmark. We release data, code, and a public leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。