针对失语症患者设计更公平的语音识别审计方法
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
- 提出三重审计漏洞应对框架,避免标准方法掩盖弱势群体问题
- 六款主流系统测试显示失语者转录错误率显著更高
- 适合关注语音识别公平性与无障碍技术的研究者和开发者
自动语音识别(ASR)系统广泛应用,亟需稳健的审计方法以确保公平的转录质量,尤其对失语症等言语障碍人群而言更为重要。尽管学术与产业界的审计已揭示不同用户群体间的性能差异,但常规审计实践常忽视关键细节,可能掩盖对边缘化群体的伤害。本文识别出三大常见审计陷阱:(1) 固守单一文本标准化方式,忽略边缘群体的偏好并掩盖性能波动;(2) 仅呈现高层次人口统计结果,未分析交叉子群体或相关声学特征下的表现差异;(3) 仅依赖词错误率(WER)这一指标,无法有效衡量生成式AI常见的幻觉等问题。为此,我们提出一个全面的审计框架,并在六款主流ASR系统上开展案例研究,发现失语者群体的表现普遍劣于对照组。我们呼吁从业者采用更严谨、以社区为导向的审计实践,以适应快速演进的ASR环境。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems' growing use warrants robust auditing approaches to ensure equitable transcription quality, especially for people with speech disorders like aphasia who disproportionately depend on ASR. While academic and industry audits have revealed performance disparities across user populations, standard auditing practices often overlook nuances that risk masking harm to marginalized groups. We identify three common pitfalls in standard ASR audits: (1) adhering to one method of text standardization, which can mask variance in ASR performance and ignore the standardization preferences of marginalized communities; (2) displaying high-level demographic findings without considering performance disparities by nuanced intersectional subgroups, or conditioning on relevant acoustic properties; and (3) reporting only one gold-standard metric (Word Error Rate), which inadequately quantifies common generative AI errors like hallucinations. We propose a holistic auditing framework addressing these pitfalls, and in a case study of six popular ASR systems, find consistently worse ASR performance for speakers with aphasia relative to a control group. We call on practitioners to implement these robust, community-driven ASR auditing practices better suited for the rapidly changing ASR landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。