构建多语言多模态音频理解新基准,推动模型跨语言跨模态推理能力评估
EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

- 设计覆盖六种语言的多模态音频评测集,融合语音、音效、音乐与图像信息
- 包含5667个选择题和13.5万条多语言翻译,揭示现有模型在跨语言场景下性能差距
- 提出轻量级微调模型Gemma3n-EXAM²,多语言与多模态任务分别提升12.4%和21.7%
近年来大规模音频语言模型(LALMs)在音频理解任务上取得显著进展,但现有评估仍主要集中于英语和单一音频领域。以往基准通常仅关注语音、音效或音乐等单模态,难以系统考察模型在多样化视觉场景下的泛化能力。本文提出EXAM²,一个覆盖六种语言、涵盖语音、音效、音乐、混合音频及视觉图像的多语言多模态音频理解基准。通过结合视觉信息与异构音频输入,EXAM²支持更真实的场景感知音频推理与跨模态理解评估。该数据集包含5,667个多项选择题、22,614张图像实例和135,684条多语言翻译。我们评估了多个开源与专有LALMs以及多模态大模型,发现其在多语言和跨模态理解上存在显著性能差距。此外,我们提出基于EXAM²训练数据微调的轻量级融合模型Gemma3n-EXAM²,在多语言设置下最高提升12.4%,在多模态评估中提升达21.7%。实验结果表明,EXAM²是一个具有挑战性的基准,将推动未来多语言多模态音频智能研究的发展。
原文摘要 · Abstract (English)
Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typically focus on a single audio modality, i.e., speech, sound, or music, limiting the systematic investigation into how these models generalize across diverse visual scenarios. In this paper, we introduce EXAM$^2$, a benchmark for multilingual and multimodal audio understanding spanning six languages and multiple modalities, including speech, sound, music, mixed-audio settings, and visual images. By incorporating visual information alongside heterogeneous audio inputs, EXAM$^2$ enables more realistic evaluation of scene-aware audio reasoning and cross-modal comprehension. EXAM$^2$ comprises $5,667$ multiple-choice questions, $22,614$ image instances, and $135,684$ multilingual translations. We evaluate state-of-the-art open-source and proprietary LALMs as well as multimodal LLMs, revealing substantial performance gaps in multilingual and cross-modal understanding. Furthermore, we propose Gemma3n-EXAM$^2$, a lightweight fusion-model fine-tuned on EXAM$^2$-train, achieves up to $12.4\%$ improvement in multilingual settings and $21.7\%$ gains in multimodal evaluation over a strong baseline. Empirical results establish EXAM$^2$ as a challenging benchmark and pioneer future multilingual and multimodal audio intelligence research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。