arXiv:2604.10014cs.CVcs.AI2026-04中稿 · ICPR 2026

评估多模态模型在不同人群和语言上的偏见,发现音频任务偏差严重。

Demographic and Linguistic Bias Evaluation in Omnimodal Language Models

  • 在文本、图像、音频、视频统一框架下评估偏见
  • 音频理解准确率低且跨年龄/性别/语言差异大
  • 适合关注AI公平性的研究人员与开发者

本文对能处理文本、图像、音频和视频的多模态语言模型在人口统计学和语言方面的偏见进行了全面评估。尽管这些模型已广泛应用,但其在不同人群和模态下的表现仍缺乏系统研究。评估了四种多模态模型在人口属性估计、身份验证、活动识别、多语言语音转录和语言识别等任务上的表现,测量了年龄、性别、肤色、语言和原籍国之间的准确率差异。结果显示,图像和视频理解任务整体表现较好,人口统计差异较小;而音频理解任务表现显著较差,存在明显的偏差,包括不同年龄组、性别和语言间的大准确率差异,以及频繁预测坍缩至少数类别。这凸显了在多模态语言模型日益应用于现实场景时,必须对其所有支持模态的公平性进行评估。

原文摘要 · Abstract (English)

This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their performance across different demographic groups and modalities is not well studied. Four omnimodal models are evaluated on tasks that include demographic attribute estimation, identity verification, activity recognition, multilingual speech transcription, and language identification. Accuracy differences are measured across age, gender, skin tone, language, and country of origin. The results show that image and video understanding tasks generally exhibit better performance with smaller demographic disparities. In contrast, audio understanding tasks exhibit significantly lower performance and substantial bias, including large accuracy differences across age groups, genders, and languages, and frequent prediction collapse toward narrow categories. These findings highlight the importance of evaluating fairness across all supported modalities as omnimodal language models are increasingly used in real-world applications.

多模态偏见评估语音识别公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。