ASR评估应承认多种转录规范,避免对言语障碍者不公平
Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation

- 提出用多种转录规范替代单一标准,避免评估偏见
- 实验证明同一语音在不同规范下WER差异可达15%以上
- 适合关注公平性、包容性的语音识别研究者
自动语音识别(ASR)评估通常以词错误率(WER)衡量系统表现,依赖单一转录参考。但这些参考由人工标注,受语义规范影响,不同规范(如逐字、非逐字、法律体)生成的转录不同,导致相同识别结果得分差异。本文指出,强制使用单一参考(参考单一体)构成认识论不公,尤其使失语症患者因语言特征被误判为错误而持续处于不利地位。评估体系缺乏解读其话语合法性的能力。我们构建解释学鸿沟理论框架,提出认识论不公距离(EID)量化指标,并基于AphasiaBank数据验证:同一语音在不同规范下,平均WER差异达15.2%。建议采用WER-Range,报告多规范下的性能范围而非预设唯一正确答案。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) evaluation compares system output to ground truth transcripts, with Word Error Rate (WER) quantifying the distance between them. But ground truth transcripts are not discovered - they are produced by human annotators following conventions that encode normative assumptions about which speech features matter. Different conventions (verbatim, non-verbatim, legal) produce different transcripts of identical speech and judge the same ASR output differently. This paper argues that reference monism - enforcing a single transcription convention as ground truth - commits epistemic injustice. Speakers with aphasia, whose speech includes clinically meaningful disfluencies, are systematically disadvantaged when evaluated against "clean" references that treat those disfluencies as errors. The harm is not merely differential performance, but that evaluative infrastructure lacks interpretive resources to recognize their contributions as legitimate. We develop a philosophical framework introducing the hermeneutical gap, formalize Epistemic Injustice Distance (EID) to measure reference monism's cost, and demonstrate empirically using AphasiaBank that WER varies depending on which convention defines ground truth. We propose WER-Range: reporting performance across legitimate conventions rather than assuming a single correct answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。