从哲学视角揭示ASR对非标准方言的系统性误识是一种深层不公。
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
- 区分中性分类与有害歧视,指出技术误识会转化成歧视
- 提出语音技术特有的三重伦理困境:时间负担、对话中断、身份关联
- 呼吁超越技术修复,承认多样语音的正当性与表达权
自动语音识别(ASR)系统已广泛介入人机交互,但对其公平性影响的研究仍显不足。本文通过哲学视角审视ASR偏见,认为对特定语音形式的系统性误识不仅是技术缺陷,更是一种对边缘化语言群体的尊重缺失,加剧了历史不公。论文区分了中性分类(discriminate1)与有害歧视(discriminate2),揭示当系统持续误识非标准方言时,前者可能演变为后者。研究识别出语音技术独有的三个伦理维度:非标准方言使用者面临的时间负担(“时间税”)、系统误识导致的对话流中断,以及语音模式与个人及文化身份的根本关联。这些因素构成不对称权力关系,现有技术公平度量无法捕捉。论文分析了ASR发展中语言标准化与多元主义之间的张力,指出当前方法常内嵌并强化有争议的语言意识形态。结论强调,解决ASR偏见需超越技术手段,须承认多样语音形式作为合法表达方式,值得技术包容。这一哲学重构为发展尊重语言多样性与说话者自主性的ASR系统提供了新路径。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems now mediate countless human-technology interactions, yet research on their fairness implications remains surprisingly limited. This paper examines ASR bias through a philosophical lens, arguing that systematic misrecognition of certain speech varieties constitutes more than a technical limitation -- it represents a form of disrespect that compounds historical injustices against marginalized linguistic communities. We distinguish between morally neutral classification (discriminate1) and harmful discrimination (discriminate2), demonstrating how ASR systems can inadvertently transform the former into the latter when they consistently misrecognize non-standard dialects. We identify three unique ethical dimensions of speech technologies that differentiate ASR bias from other algorithmic fairness concerns: the temporal burden placed on speakers of non-standard varieties ("temporal taxation"), the disruption of conversational flow when systems misrecognize speech, and the fundamental connection between speech patterns and personal/cultural identity. These factors create asymmetric power relationships that existing technical fairness metrics fail to capture. The paper analyzes the tension between linguistic standardization and pluralism in ASR development, arguing that current approaches often embed and reinforce problematic language ideologies. We conclude that addressing ASR bias requires more than technical interventions; it demands recognition of diverse speech varieties as legitimate forms of expression worthy of technological accommodation. This philosophical reframing offers new pathways for developing ASR systems that respect linguistic diversity and speaker autonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。