arXiv:2510.04584cs.CLcs.SD2025-10中稿 · Interspeech 2026被引 5

音频大模型在选择题评估中对选项顺序和问题表述敏感,需改进评测方法。

Robustness assessment of large audio language models in multiple-choice evaluation

  • 分析三种基准与四类模型在选择题中的表现差异。
  • 发现选项顺序和问题表述变化会显著影响模型准确率。
  • 提出更鲁棒的评测协议,提升评估细致度与可靠性。

近年来,大型音频语言模型(LALMs)主要通过多项选择题问答(MCQA)框架进行评估。然而,细微变化如选项顺序调整会导致结果显著不同。现有MCQA框架未考虑这种变异性,仅报告单一准确率。本文系统研究了三个基准(MMAU、MMAR、MMSU)及四种模型(Audio Flamingo 2、Audio Flamingo 3、Qwen2.5-Omni-7B-Instruct、Kimi-Audio-7B-Instruct)的表现。结果表明,模型不仅对选项顺序敏感,也对问题和选项的改写形式敏感。为此,我们提出一种更简洁的评估协议与指标,能有效捕捉细微变化,提供更详尽的模型评估报告。

原文摘要 · Abstract (English)

Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.

音频模型评测方法多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。