首个系统评估多语言语音大模型偏见的研究,发现语言和选项顺序影响更大。
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
- 构建跨语言语音数据集,测试不同语言、口音、性别和选项顺序下的模型表现
- 模型对语言和选项顺序敏感,但对性别和口音相对鲁棒,结构偏差被语音放大
- 提供可复现的评估框架,适合关注语音模型公平性的研究者
本研究首次系统性地考察多语言多模态大模型中的语音偏见。我们构建并发布了BiasInEar数据集,基于Global MMLU Lite,涵盖英语、中文和韩语,按性别和口音平衡,共包含70.8小时(约4,249分钟)语音与11,200个问题。采用准确率、熵、APES和Fleiss' κ四项互补指标,在语言、口音、性别和选项顺序扰动下评估九个代表性模型。结果表明,多模态大模型对人口统计因素相对鲁棒,但对语言和选项顺序高度敏感,提示语音可能放大现有结构偏差。此外,模型架构与推理策略显著影响跨语言鲁棒性。本研究建立统一评估框架,弥合文本与语音评估之间的差距。资源详见https://github.com/ntunlplab/BiasInEar。
原文摘要 · Abstract (English)
This work presents the first systematic investigation of speech bias in multilingual MLLMs. We construct and release the BiasInEar dataset, a speech-augmented benchmark based on Global MMLU Lite, spanning English, Chinese, and Korean, balanced by gender and accent, and totaling 70.8 hours ($\approx$4,249 minutes) of speech with 11,200 questions. Using four complementary metrics (accuracy, entropy, APES, and Fleiss' $κ$), we evaluate nine representative models under linguistic (language and accent), demographic (gender), and structural (option order) perturbations. Our findings reveal that MLLMs are relatively robust to demographic factors but highly sensitive to language and option order, suggesting that speech can amplify existing structural biases. Moreover, architectural design and reasoning strategy substantially affect robustness across languages. Overall, this study establishes a unified framework for assessing fairness and robustness in speech-integrated LLMs, bridging the gap between text- and speech-based evaluation. The resources can be found at https://github.com/ntunlplab/BiasInEar.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。