对比13种去偏方法,找出多标签语音情感识别中兼顾公平与准确的最佳方案。
EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
- 系统性比较13种去偏技术,涵盖预处理到分布鲁棒优化
- 在性别不平衡数据下,部分方法可降低性能差距20%以上
- 适合关注模型公平性的语音识别研究者和工程师
语音情感识别(SER)系统常存在性别偏差。然而,现有去偏方法在多标签场景下的有效性与鲁棒性仍缺乏深入研究。为填补这一空白,我们提出EMO-Debias,对13种去偏方法在多标签SER中的表现进行大规模对比。研究涵盖预处理、正则化、对抗学习、有偏学习及分布鲁棒优化等策略。实验基于真人演绎与自然语境情感数据集,使用WavLM和XLSR表征,在性别失衡条件下评估各方法表现。分析量化了公平性与准确率之间的权衡,识别出能在不牺牲整体性能前提下持续缩小性别性能差距的方法。研究结果为选择有效去偏策略提供了可操作洞见,并揭示了数据分布的影响。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) systems often exhibit gender bias. However, the effectiveness and robustness of existing debiasing methods in such multi-label scenarios remain underexplored. To address this gap, we present EMO-Debias, a large-scale comparison of 13 debiasing methods applied to multi-label SER. Our study encompasses techniques from pre-processing, regularization, adversarial learning, biased learners, and distributionally robust optimization. Experiments conducted on acted and naturalistic emotion datasets, using WavLM and XLSR representations, evaluate each method under conditions of gender imbalance. Our analysis quantifies the trade-offs between fairness and accuracy, identifying which approaches consistently reduce gender performance gaps without compromising overall model performance. The findings provide actionable insights for selecting effective debiasing strategies and highlight the impact of dataset distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。