arXiv:2510.00628cs.SDcs.CL2025-10中稿 · Interspeech 2026被引 6

发现大模型回答顺序影响判断,可能误导评估结果

Hearing the Order: Investigating Position Bias in Large Audio-Language Models

  • 测试六款大音频模型在三种评测集上的顺序敏感性
  • 答案顺序打乱后准确率波动最高达24%,排名可能逆转
  • 建议用打乱顺序策略降低偏差,适合评估研究者参考

大型音频语言模型(LALMs)常用于涉及有序选项推理的任务。当前一个未解问题是:模型预测是否受答案选项顺序影响,这将体现位置偏差并削弱其可靠性。本文首次系统性地考察了这一问题。通过在三个主流基准及其语音版本上对六款LALMs进行广泛实验,我们证明无一模型能免疫此偏差。打乱答案顺序可导致性能波动高达24%,甚至改变模型排名,引发对现有评估方法可靠性的担忧。我们进一步研究基于排列的策略,发现其在多数情况下可缓解偏差。本工作旨在提高对此问题的认识,并推动该方向的后续研究。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) are often used in tasks that involve reasoning over ordered options. An open question is whether their predictions are influenced by the order of answer choices, which would indicate a form of position bias and undermine their reliability. In this paper, we identify and analyze this problem in LALMs. We demonstrate that no model is immune to this bias through extensive experiments on six LALMs across three widely used benchmarks and their spoken counterparts. Shuffling the order of answer options can cause performance fluctuations of up to 24% and even change model rankings, raising concerns about the reliability of current evaluation practices. We also study permutation-based strategies and show that they can mitigate bias in most cases. Our work represents the first systematic investigation of this issue in LALMs, and we hope it raises awareness and motivates further research in this direction.

音频模型位置偏差评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。