分析方言语音差异如何导致主流语音识别系统对非裔说话人更不准确。
A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
- 基于语音社会学标注,量化不同族裔发音差异带来的识别错误。
- 非裔说话人因低后元音合并抵抗等特征,错误率显著更高。
- 研究可帮助改进训练数据,减少语音识别中的种族偏见。
本文利用太平洋西北英语语料库(PNWE),系统评估了四种主流商业自动语音识别(ASR)系统中的种族偏见。分析涵盖四个族裔背景(非裔、白人、奇卡诺/奇卡娜、亚卡马)说话人的转录准确性,并考察社会语音变异如何影响系统性能差异。提出一种启发式确定的语音错误率(PER)指标,将识别错误与源自社会语音标注的特定语言学变量关联。对十一项社会语音特征的分析表明,元音质量变化,特别是对低后元音合并和前鼻音合并模式的抵抗,与不同族裔间系统错误率差异系统相关,其中非裔说话人各系统中均表现最差。研究证实,方言语音变异的声学建模而非词汇或句法因素,仍是商业ASR系统偏见的主要来源。该研究确立了PNWE语料库在语音技术偏见评估中的价值,并为通过针对性地纳入社会语音多样性来提升ASR性能提供可操作建议。
原文摘要 · Abstract (English)
This paper presents a systematic evaluation of racial bias in four major commercial automatic speech recognition (ASR) systems using the Pacific Northwest English (PNWE) corpus. We analyze transcription accuracy across speakers from four ethnic backgrounds (African American, Caucasian American, ChicanX, and Yakama) and examine how sociophonetic variation contributes to differential system performance. We introduce a heuristically-determined Phonetic Error Rate (PER) metric that links recognition errors to specific linguistically motivated variables derived from sociophonetic annotation. Our analysis of eleven sociophonetic features reveals that vowel quality variation, particularly resistance to the low-back merger and pre-nasal merger patterns, is systematically associated with differential error rates across ethnic groups, with the most pronounced effects for African American speakers across all evaluated systems. These findings demonstrate that acoustic modeling of dialectal phonetic variation, rather than lexical or syntactic factors, remains a primary source of bias in commercial ASR systems. The study establishes the PNWE corpus as a valuable resource for bias evaluation in speech technologies and provides actionable guidance for improving ASR performance through targeted representation of sociophonetic diversity in training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。