arXiv:2603.21050cs.SD2026-03中稿 · presentation at IN…被引 3

针对多语言语音大模型的情感识别性别偏见,提出新方法降低偏差并提升性能。

ERM-MinMaxGAP: Benchmarking and Mitigating Gender Bias in Multilingual Multimodal Speech-LLM Emotion Recognition

  • 设计跨语言多模态基准,量化不同语言中性别表现差异。
  • 在英语、日语、德语上,性能提升5.5%~5.0%,性别差距缩小0.1%~1.4%。
  • 创新性引入自适应权重与最大损失差正则化,提升公平性与泛化能力。

语音情感识别(SER)系统可能存在性别相关性能差异,但多语言语音大模型在不同语言和模态下的偏见表现尚不明确。本文基于MELD-ST构建了一个涵盖英语、日语和德语的多语言多模态基准,用于量化语言特异性情感识别性能及性别差距。研究发现偏见具有强语言依赖性,且多模态融合并不能可靠提升公平性。为此,提出ERM-MinMaxGAP,一种面向公平性的训练目标,在经验风险最小化基础上引入自适应公平性权重机制与新的MinMaxGAP正则项,以最小化每种语言和模态下的最大男女损失差距。基于Qwen2-Audio骨干网络,该方法在单模态和多模态设置下分别提升性能5.5%和5.0%,同时将整体性别偏差差距降低0.1%和1.4%。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) systems can exhibit gender-related performance disparities, but how such bias manifests in multilingual speech LLMs across languages and modalities is unclear. We introduce a novel multilingual, multimodal benchmark built on MELD-ST, spanning English, Japanese, and German, to quantify language-specific SER performance and gender gaps. We find bias is strongly language-dependent, and multimodal fusion does not reliably improve fairness. To address these, we propose ERM-MinMaxGAP, a fairness-informed training objective, which augments empirical risk minimization (ERM) with a proposed adaptive fairness weight mechanism and a novel MinMaxGAP regularizer on the maximum male-female loss gap within each language and modality. Building upon the Qwen2-Audio backbone, our ERM-MinMaxGAP approach improves multilingual SER performance by 5.5% and 5.0% while reducing the overall gender bias gap by 0.1% and 1.4% in the unimodal and multimodal settings, respectively.

语音识别多模态公平性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。