用新指标揭示语音识别中对少数群体的系统性偏见
Beyond Word Error Rate: Auditing the Diversity Tax in Speech Recognition through Dataset Cartography
- 引入样本难度指数SDI,量化声学与人口因素如何导致模型失效
- 发现EmbER和SemDist能暴露WER忽略的语义偏差和模型分歧
- 为部署前评估语音识别公平性提供可操作的审计框架
自动语音识别(ASR)系统主要依赖词错误率(WER)进行评估,但单纯基于词级统计的指标无法捕捉语义保真度,常掩盖‘多样性税’——即边缘化和非典型说话者因系统性识别失败而承受的不公负担。本文通过系统性评估非线性与语义类指标,揭示仅依赖词汇计数的局限性。为实现严格模型审计,提出样本难度指数(SDI),量化内在人口与声学因素对模型失败的影响。通过数据制图法映射SDI,发现EmbER与SemDist能揭示韦尔忽略的系统性偏见及模型间分歧。研究成果标志着构建稳健审计框架的初步进展,助力开发者在部署前识别并缓解ASR差异。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems are predominantly evaluated using the Word Error Rate (WER). However, raw token-level metrics fail to capture semantic fidelity and routinely obscures the `diversity tax', the disproportionate burden on marginalized and atypical speaker due to systematic recognition failures. In this paper, we explore the limitations of relying solely on lexical counts by systematically evaluating a broader class of non-linear and semantic metrics. To enable rigorous model auditing, we introduce the sample difficulty index (SDI), a novel metric that quantifies how intrinsic demographic and acoustic factors drive model failure. By mapping SDI on data cartography, we demonstrate that metrics EmbER and SemDist expose hidden systemic biases and inter-model disagreements that WER ignores. Finally, our findings are the first steps towards a robust audit framework for prospective safety analysis, empowering developers to audit and mitigate ASR disparities prior to deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。