arXiv:2603.24549cs.CLcs.AI2026-03被引 2

分析纽卡斯尔方言中语音识别偏见,发现错误与社会因素相关

A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English

  • 基于3000+转录错误,从语言学角度分析方言差异导致的识别失败
  • 男性及两头年龄组错误率更高,方言特征如元音质量影响识别准确率
  • 呼吁将方言数据和社区参与纳入语音技术开发以减少社会偏见

自动语音识别(ASR)系统广泛应用于日常沟通、教育、医疗和工业领域,但其在不同说话人中的表现存在差异,尤其当方言与训练数据中的主流口音偏离时。本研究通过社会语言学视角,分析英国东北部纽卡斯尔方言对当前先进商业ASR系统的影响。使用迪亚克罗尼电子泰恩赛德英语语料库(DECTE)中的自发口语数据,评估了系统输出并细致分析了超过3000个转录错误。错误按语言领域分类,并关联性别、年龄和经济社会地位等社会变量进行考察。此外,针对特定元音特征的声学案例研究显示,渐进性语音变异会直接导致误识别。结果表明,多数错误源于音系变异,反复出现的问题与方言特有的元音质量、喉塞音化、本地词汇及非标准语法形式有关。错误率在社会群体间呈现差异,男性及年龄两极群体错误频率更高。研究证明ASR错误并非随机,而是具有社会模式,可从社会语言学角度解释。因此,该研究强调在语音技术评估与开发中引入社会语言学专业知识的重要性,指出更公平的ASR系统需明确关注方言差异和基于社区的语音数据。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems are widely used in everyday communication, education, healthcare, and industry, yet their performance remains uneven across speakers, particularly when dialectal variation diverges from the mainstream accents represented in training data. This study investigates ASR bias through a sociolinguistic analysis of Newcastle English, a regional variety of North-East England that has been shown to challenge current speech recognition technologies. Using spontaneous speech from the Diachronic Electronic Corpus of Tyneside English (DECTE), we evaluate the output of a state-of-the-art commercial ASR system and conduct a fine-grained analysis of more than 3,000 transcription errors. Errors are classified by linguistic domain and examined in relation to social variables including gender, age, and socioeconomic status. In addition, an acoustic case study of selected vowel features demonstrates how gradient phonetic variation contributes directly to misrecognition. The results show that phonological variation accounts for the majority of errors, with recurrent failures linked to dialect-specific features like vowel quality and glottalisation, as well as local vocabulary and non-standard grammatical forms. Error rates also vary across social groups, with higher error frequencies observed for men and for speakers at the extremes of the age spectrum. These findings indicate that ASR errors are not random but socially patterned and can be explained from a sociolinguistic perspective. Thus, the study demonstrates the importance of incorporating sociolinguistic expertise into the evaluation and development of speech technologies and argues that more equitable ASR systems require explicit attention to dialectal variation and community-based speech data.

语音识别方言偏见社会语言学公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。