破解语音识别中的语言殖民主义,让边缘语言声音被机器听懂
Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
- 用三重伤害框架识别语音系统对少数语言的歧视性误识
- 提出七层语境模型,量化语言多样性在语音系统中的缺失程度
- 倡导社区共治,让原住民群体参与语音系统设计与评估
本文聚焦自动语音识别(ASR)及其驱动的语音交互界面在公共服务、医疗和教育中的应用。我们指出,低资源、原住民及非标准语言变体在语音识别中的持续失败不仅是技术问题,更是隐含的言语政策,延续了殖民语言等级制度。基于语言资本、种族语言意识形态、语言政策研究和去殖民计算理论,本文揭示数据、指标与模型先验如何决定哪些声音能被机器识别。提出‘三重伤害’(3M)分类法——误识、错位、不信任,并构建七层语境化模型以刻画语音识别中的语言多样性状况。进一步提出参与式框架与最小审计协议,将受影响社区定位为共同设计者、评估者与治理伙伴。
原文摘要 · Abstract (English)
This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy---Misrecognition, Misalignment, and Mistrust---and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。