评估语音识别中基于音素的模型在不同人群中的偏差,发现即使考虑音素相似性仍存在性能差距。
Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

- 对比两个开源音素识别系统在多种口音和语言上的表现
- 发现性别、年龄、族裔等群体间仍有显著识别误差差异
- 提出软音素错误率度量,更合理评估音素相近的替换
自动语音识别(ASR)系统的普及引发了对种族、年龄、性别和口音等人口统计偏差的关注,这些偏差通常源于训练数据不平衡。现有研究多聚焦于基于字符的标准ASR系统,而对生成国际音标(IPA)表示的音素基系统关注较少。随着ASR向多语言支持和低资源语言建模发展,基于IPA的层成为关键的语言无关基础。本研究评估了两种最先进的开源ASR系统WhisperIPA和ZIPA在多样口音和语言来源下的表现,使用现有多语言语料库及带有社会人口标注的英语语料库进行测试。通过标准音素错误率(PER)和一种新提出的允许语言学上相似音素替换的软音素错误率(Soft PER)进行比较。分析显示,尽管考虑了可接受的音素变异,不同语言和人口群体(如性别、口音、族裔、年龄)间的性能差异依然存在。这些发现揭示了潜在偏见来源,为构建更包容、更具语言鲁棒性的音素基ASR系统提供依据。代码与数据将公开共享。
原文摘要 · Abstract (English)
The popularization of automatic speech recognition (ASR) systems has increased exploration of the demographic biases related to race, age, gender, and accent, often formed from imbalanced training data. Most of these studies focused on standard grapheme-based ASR systems with comparatively little emphasis on phoneme-based systems, such as models that produce International Phonetic Alphabet (IPA) representations. As ASR systems shift toward multilingual support and low-resource language modeling, IPA-based layers serve as a critical, language-agnostic foundation. In this study, we evaluate the performance of two state-of-the-art open-source ASR systems, WhisperIPA and ZIPA, that generate IPA transcriptions across diverse accents and language sources. Our evaluation includes existing multilingual speech corpora and demographically annotated English-language corpora. We measure model performance by comparing model-generated IPA transcriptions against grapheme-to-phoneme (G2P) systems using both standard phoneme error rate (PER) and a proposed Soft PER metric that tolerates linguistically similar phoneme substitutions. Our analysis examines how performance varies across languages and demographic groups such as gender, accent, ethnicity, and age, revealing persistent disparities even after accounting for acceptable phonemic variation. These findings provide insight into potential sources of bias and inform the development of more inclusive and linguistically robust phoneme-based ASR systems. Our code and data will be made publicly available to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。