arXiv:2605.10615cs.CL2026-05

提出ASR公平性评估新标准,避免误判歧视群体。

Responsible Benchmarking of Fairness for Automatic Speech Recognition

论文配图:Responsible Benchmarking of Fairness for Automatic Speech Recognition
图 1 · 摘自论文原文
  • 按假设定制公平性度量,避免一刀切评估
  • 发现单一群体划分会掩盖真实歧视对象
  • 建议细粒度分析多维人口变量交叉影响

许多研究显示自动语音识别(ASR)系统在不同说话人群体(SG)间表现不均。然而,现有研究得出该结论的方法并不一致。为推动未来研究的可靠性,本文基于机器学习公平性、社会科学和语音科学的文献,提出评估ASR公平性的最佳实践。首先强调需精准定义所检验的公平性假设,并相应设计度量指标。随后分析多个用于评估ASR公平性的基准,指出若不严谨考察群体间的交叉关系,其结果易被误解。研究发现,仅基于单一异质群体(如现有公平性基准所定义)评估公平性,可能导致错误识别出实际被系统歧视的群体。因此,应尽可能细致地分析公平性语料库中所有可用的人口学变量交叉作用,以揭示虚假相关性。

原文摘要 · Abstract (English)

Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which such studies arrive at this conclusion is inconsistent. To pave the wayfor more reliable results in future studies, we lay out best practices for benchmarking ASR fairness based on literaturefrom machine learning fairness, social sciences, and speech science. We first describe the importance of preciselythe fairness hypothesis being interrogated, and tailoring fairness metrics to apply specifically to said hypothesis.We then examine several benchmarks used to rate ASR systems on fairness and discuss how their results can bemisconstrued without assiduous oversight into the intersections between SG's. We find that evaluating fairnessbased on single heterogeneous SG's, such as they are defined in fairness benchmarks, can lead to misidentifyingwhich SG's are actually being mistreated by ASR systems. We advocate for as fine-grained an analysis as possibleof the intersectionality of as many demographic variables as are available in the metadata of fairness corpora in orderto tease out such spurious correlations

ASR公平性交叉性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。