分析荷兰语语音识别模型在不同性别间的性能偏差,揭示系统性不公平问题。
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
- 基于荷兰语数据集,对比不同性别语音的识别误差
- 所有模型规模下女性语音的词错误率显著更高
- 为自动字幕等应用提供公平性评估参考
近期研究显示,如Whisper等先进语音识别系统常对不同人群表现出预测偏差。本研究聚焦Whisper模型在荷兰语语音数据上的表现差异,数据来自Common Voice及荷兰国家广播机构。通过词错误率(WER)、字符错误率及基于BERT的语义相似度,分析了不同性别群体的表现。采用Weerts等人(2022)的道德框架评估服务质量损害与公平性,探讨这些偏差对自动字幕等应用的影响。统计检验表明,所有模型规模下性别间均存在显著的词错误率差异,揭示出系统性偏差。
原文摘要 · Abstract (English)
Recent research has shown that state-of-the-art (SotA) Automatic Speech Recognition (ASR) systems, such as Whisper, often exhibit predictive biases that disproportionately affect various demographic groups. This study focuses on identifying the performance disparities of Whisper models on Dutch speech data from the Common Voice dataset and the Dutch National Public Broadcasting organisation. We analyzed the word error rate, character error rate and a BERT-based semantic similarity across gender groups. We used the moral framework of Weerts et al. (2022) to assess quality of service harms and fairness, and to provide a nuanced discussion on the implications of these biases, particularly for automatic subtitling. Our findings reveal substantial disparities in word error rate (WER) among gender groups across all model sizes, with bias identified through statistical testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。