提出新公平性度量,揭示语音情感识别中的性别偏见
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias

- 通过联合建模身份属性与模型误差,捕捉分配偏见
- 在CREMA-D数据集上发现HuBERT和WavLM存在性别偏见
- 可量化各身份属性对偏见的绝对贡献,适合伦理审查
语音情感识别(SER)系统在心理健康和教育等敏感领域应用日益广泛,但其偏差预测可能造成伤害。传统公平性度量如平等机会和人口均等性常忽略身份属性与模型预测之间的联合依赖关系。本文提出一种面向SER的公平性建模方法,通过学习身份属性与模型误差的联合关系,显式捕捉分配偏见。我们在合成数据上验证了该公平性度量,随后将其应用于在CREMA-D数据集上微调的HuBERT与WavLM模型。结果表明,所提方法能更准确捕获受保护属性与偏见间的互信息,并量化单个属性对基于自监督学习的SER模型偏见的绝对贡献。此外,分析显示HuBERT与WavLM均存在性别偏见迹象。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) systems have growing applications in sensitive domains such as mental health and education, where biased predictions can cause harm. Traditional fairness metrics, such as Equalised Odds and Demographic Parity, often overlook the joint dependency between demographic attributes and model predictions. We propose a fairness modelling approach for SER that explicitly captures allocative bias by learning the joint relationship between demographic attributes and model error. We validate our fairness metric on synthetic data, then apply it to evaluate HuBERT and WavLM models finetuned on the CREMA-D dataset. Our results indicate that the proposed fairness model captures more mutual information between protected attributes and biases and quantifies the absolute contribution of individual attributes to bias in SSL-based SER models. Additionally, our analysis reveals indications of gender bias in both HuBERT and WavLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。