arXiv:2505.11572cs.SDcs.CL2025-05中稿 · INTERSPEECH 2025被引 5

构建语音识别公平性评估基准,揭示主流模型的群体差异。

ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems

  • 用混合效应泊松回归分析不同群体表现差异。
  • 发现顶尖模型在不同人群间误差率差距显著。
  • 适合关注AI公平性的研究者与开发者参考。

自动语音识别(ASR)系统已广泛应用于日常场景,但不同人口群体间的性能差异仍显著存在。本文提出ASR-FAIRBENCH排行榜,用于实时评估ASR模型的准确性和公平性。基于Meta的Fair-Speech数据集,该数据集涵盖多样化的人口特征,我们采用混合效应泊松回归模型计算整体公平性得分,并结合传统指标词错误率(WER),构建公平性调整后的ASR得分(FAAS),形成综合评估框架。结果揭示了当前最先进的ASR模型在不同群体间存在显著性能差异,为推动更包容的ASR技术发展提供了基准。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems have become ubiquitous in everyday applications, yet significant disparities in performance across diverse demographic groups persist. In this work, we introduce the ASR-FAIRBENCH leaderboard which is designed to assess both the accuracy and equity of ASR models in real-time. Leveraging the Meta's Fair-Speech dataset, which captures diverse demographic characteristics, we employ a mixed-effects Poisson regression model to derive an overall fairness score. This score is integrated with traditional metrics like Word Error Rate (WER) to compute the Fairness Adjusted ASR Score (FAAS), providing a comprehensive evaluation framework. Our approach reveals significant performance disparities in SOTA ASR models across demographic groups and offers a benchmark to drive the development of more inclusive ASR technologies.

语音识别公平性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。