MAEB benchmark评估50+模型在30项跨语言音频任务中的表现,揭示音视频模型的性能短板。
MAEB: Massive Audio Embedding Benchmark
- 构建覆盖100+语言的30项音频任务基准,涵盖语音、音乐与跨模态推理。
- 无模型在所有任务上领先:声学理解强的模型在语言任务中表现差,反之亦然。
- 结果与音频大模型性能高度相关,适合研究多模态模型泛化能力者参考。
我们提出大规模音频嵌入基准(MAEB),涵盖30个任务,覆盖语音、音乐、环境声音及跨模态音频-文本推理,涉及100多种语言。评估了50多个模型,发现无单一模型在所有任务上占优:对比型音频-文本模型在环境声音分类(如ESC50)中表现优异,但在多语言语音任务(如SIB-FLEURS)中近乎随机;而语音预训练模型则呈现相反趋势。聚类任务对所有模型仍具挑战,最优模型仅达有限效果。我们观察到,在声学理解方面表现好的模型往往在语言任务中表现差,反之亦然。此外,音频编码器在MAEB上的表现与在音频大语言模型中的表现高度相关。MAEB源自包含98个任务的MAEB+,在保持任务多样性的同时降低评估成本,并集成至MTEB生态,实现文本、图像与音频的统一评估。我们已开源MAEB及全部98个任务的代码、数据与排行榜,详见https://github.com/embeddings-benchmark/mteb。
原文摘要 · Abstract (English)
We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 50+ models and find that no single model dominates across all tasks: contrastive audio-text models excel at environmental sound classification (e.g., ESC50) but score near random on multilingual speech tasks (e.g., SIB-FLEURS), while speech-pretrained models show the opposite pattern. Clustering remains challenging for all models, with even the best-performing model achieving only modest results. We observe that models excelling on acoustic understanding often perform poorly on linguistic tasks, and vice versa. We also show that the performance of audio encoders on MAEB correlates highly with their performance when used in audio large language models. MAEB is derived from MAEB+, a collection of 98 tasks. MAEB is designed to maintain task diversity while reducing evaluation cost, and it integrates into the MTEB ecosystem for unified evaluation across text, image, and audio modalities. We release MAEB and all 98 tasks along with code and a leaderboard at https://github.com/embeddings-benchmark/mteb.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。