arXiv:2505.14449eess.AScs.CL2025-05中稿 · InterSpeech 2025被引 8

通过伪标签与无监督学习,缓解语音情绪识别中的群体偏差

Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach

  • 用预训练模型生成伪标签,结合k-means聚类实现隐式群体推断
  • 无监督方法使公平性指标提升超4.6%,准确率下降不足3.6%
  • 无需显式人口信息,有效降低种族与年龄带来的识别偏差

尽管群体差异和性能偏倚在计算研究中日益受到关注,但分类型语音情绪识别(SER)中的公平性仍研究不足。现有方法通常依赖显式的人口统计标签,但因隐私问题难以获取。为此,我们提出隐式人口推断(IDI)模块,利用预训练模型生成的伪标签,并通过k-means聚类实现无监督学习,以缓解SER中的偏差。实验表明,伪标签驱动的IDI将群体差异减少,公平性指标提升超过28%,且SER准确率下降不足2%。无监督IDI进一步使公平性指标提升超4.6%,准确率下降低于3.6%。深入分析显示,该方法能持续缓解种族与年龄差异,适用于缺乏显式人口信息的场景。

原文摘要 · Abstract (English)

While subgroup disparities and performance bias are increasingly studied in computational research, fairness in categorical Speech Emotion Recognition (SER) remains underexplored. Existing methods often rely on explicit demographic labels, which are difficult to obtain due to privacy concerns. To address this limitation, we introduce an Implicit Demography Inference (IDI) module that leverages pseudo-labeling from a pre-trained model and unsupervised learning using k-means clustering to mitigate bias in SER. Our experiments show that pseudo-labeling IDI reduces subgroup disparities, improving fairness metrics by over 28% with less than a 2% decrease in SER accuracy. Also, the unsupervised IDI yields more than a 4.6% improvement in fairness metrics with a drop of less than 3.6% in SER performance. Further analyses reveal that the unsupervised IDI consistently mitigates race and age disparities, demonstrating its potential when explicit demographic information is unavailable.

语音情绪识别公平性无监督学习伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。