研究语音识别中性别偏差,发现训练数据比例影响识别准确率
Exploring Gender Disparities in Automatic Speech Recognition Technology
- 用LibriSpeech和Whisper小模型分析不同性别数据比例对性能影响
- 最优公平性出现在特定性别分布,而非50-50均衡
- 音高变化是影响识别准确率的关键因素,适合关注公平性的研究者
本研究探讨自动语音识别(ASR)系统在性别维度上的公平性与性能影响因素,超越传统的人口统计学分析。基于LibriSpeech数据集与Whisper小模型,分析训练数据中不同性别表示对性能的影响。研究发现,训练数据中的性别比例与ASR性能之间存在复杂关系,最佳公平性出现在特定性别分布,而非简单的50-50平衡。此外,音高变化等声学特征显著影响识别准确率。该研究深化了对ASR系统偏见的理解,强调精心构建训练数据对缓解性别偏差的重要性。
原文摘要 · Abstract (English)
This study investigates factors influencing Automatic Speech Recognition (ASR) systems' fairness and performance across genders, beyond the conventional examination of demographics. Using the LibriSpeech dataset and the Whisper small model, we analyze how performance varies across different gender representations in training data. Our findings suggest a complex interplay between the gender ratio in training data and ASR performance. Optimal fairness occurs at specific gender distributions rather than a simple 50-50 split. Furthermore, our findings suggest that factors like pitch variability can significantly affect ASR accuracy. This research contributes to a deeper understanding of biases in ASR systems, highlighting the importance of carefully curated training data in mitigating gender bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。