让语音识别模型更公平,减少不同人群间的识别差异。
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
- 用反向梯度层抑制与群体身份相关的特征
- 在多人口统计数据集上显著降低群体间识别差距
- 适合关注语音系统公平性的研究与开发者
大规模语音识别模型在准确性和鲁棒性方面取得了显著进展,但在实际应用中,公平性问题仍被忽视。本文提出 FairASR,通过学习对群体身份不敏感的表示,实现跨人口统计群体的公平泛化。基于多人口统计数据集,该方法利用梯度反转层抑制与群体相关的判别特征,同时通过无监督对比损失保留通用语音模式。实验表明,FairASR 在保持竞争力整体性能的同时,显著降低了不同群体间的识别性能差异。
原文摘要 · Abstract (English)
Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-world applications. In this work, we introduce FairASR, a system that mitigates demographic bias by learning representations that are uninformative about group membership, enabling fair generalization across demographic groups. Leveraging a multi-demographic dataset, our approach employs a gradient reversal layer to suppress demographic-discriminative features while maintaining the ability to capture generalizable speech patterns through an unsupervised contrastive loss. Experimental results show that FairASR delivers competitive overall ASR performance while significantly reducing performance disparities across different demographic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。