arXiv:2605.30472cs.CL2026-05

多模态语音模型在不同人脸下识别准确率差异高达4.05,暴露了潜在偏见。

Your Multimodal Speech Model Says I Have a Face for Radio

  • 用相同音频配不同人脸生成视频,测试多模态语音识别偏差
  • 多款模型在性别、种族交叉维度上误差最高上升4.05个点
  • 提示开发者警惕模态增加带来的新偏见,需主动评估与沟通

随着大模型在语言任务上的进步,研究者正构建处理更多模态数据的多模态与全模态模型。例如,语音识别模型已扩展至音视频数据,用于降噪和多模态字幕生成。尽管单模态下的性能与偏见已有广泛研究,但新增模态如何影响这些因素尚不明确,而人类对多模态信号也存在偏见。因此,我们首次提出多模态语音识别的偏见评估方法:创建将不同人脸与相同音频配对的视频,测量语音转录准确性的变化。结果发现,mWhisper-Flamingo 和 Gemini 模型在自我声明的性别、种族及其交叉维度上存在显著服务质量差异,词错误率(WER)最高上升达4.05点。研究指出,开发者亟需评估、修复并公开此类局限性,因为增加模态信号并不必然提升性能,反而可能引发偏见。

原文摘要 · Abstract (English)

As large neural models have become better at language tasks, researchers are increasingly building multi- and omnimodal models that handle more modalities of data. One example is the expansion of speech recognition models to audio-visual data for noise mitigation and multimodal subtitling. While performance and bias have been studied extensively in the single-modality regime, it is unknown how new modalities affect this, even though they produce biases in humans. We therefore propose the first bias evaluation of multimodal speech recognition, where we create videos pairing different faces with the same audio, and measure changes in speech transcription accuracy. We find large quality-of-service differences across mWhisper-Flamingo and Gemini models, with drops of up to 4.05 word error rate points, across self-declared gender, ethnicity, and their intersection. Our findings point to a priority for developers to evaluate, fix, and communicate such limitations, as providing more signals through additional modalities is not necessarily better, and may even lead to biased outcomes.

多模态语音识别偏见检测模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。