不同母语者在英语语音识别中的表现差异,与母语距离英语远近有关。
Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition Systems

- 基于母语与英语的语系距离,分析语音识别误差率变化规律
- 跨数据集和模型验证,母语距离与识别错误率显著相关(p<0.001)
- 深层声学层中存在按母语分组的潜在表示空间结构
尽管近年来自动语音识别(ASR)模型取得了显著进展,但在不同说话人群体间仍存在性能差异,尤其体现在母语(L1)与英语语系距离较远的说话人。本文研究了母语背景与英语语音识别性能之间的关系。通过实证分析发现,说话人母语与英语的距离与其语音识别错误率之间存在系统性关联,且该关联在不同数据集和模型间强度不一。在考虑数据集层面变异的Tweedie混合效应模型中,这一关联在所有评估模型中均具有统计显著性(p<0.001)。此外,对潜在空间的分析显示,在多数评估架构的深层声学层中,存在基于母语的表征空间分离现象。
原文摘要 · Abstract (English)
While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。