arXiv:2409.03059cs.CL2024-09被引 3

分析非裔美国人英语转录中的风格差异对语音识别评测的影响

Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

  • 对比6种转录版本,聚焦原样性与非裔英语语法特征
  • 发现转录风格差异显著影响词错误率的可比性
  • 揭示训练数据人工转录者的决策如何塑造语音识别结果

评估自动语音识别(ASR)系统与人工转录表现的常用准确率指标,会混淆多种误差来源。当训练集与测试集存在风格差异时,如原样性与非原样性转录方式的不同,将显著影响ASR性能评估。这一问题在代表性不足的语言变体中尤为突出,其语音到文字的映射标准不统一。本文对10小时非裔美国人英语(AAE)语音的6种转录版本(4种人工、2种ASR生成)进行分类,聚焦原样性特征与AAE句法形态特征,研究这些类别如何影响通过词错误率(WER)比较转录质量的效果。结果表明,转录风格差异显著干扰了不同转录间的可比性,且整体分析揭示了ASR输出本质上是训练数据中人工转录者决策的产物。

原文摘要 · Abstract (English)

Common measures of accuracy used to assess the performance of automatic speech recognition (ASR) systems, as well as human transcribers, conflate multiple sources of error. Stylistic differences, such as verbatim vs non-verbatim, can play a significant role in ASR performance evaluation when differences exist between training and test datasets. The problem is compounded for speech from underrepresented varieties, where the speech to orthography mapping is not as standardized. We categorize the kinds of stylistic differences between 6 transcription versions, 4 human- and 2 ASR-produced, of 10 hours of African American English (AAE) speech. Focusing on verbatim features and AAE morphosyntactic features, we investigate the interactions of these categories with how well transcripts can be compared via word error rate (WER). The results, and overall analysis, help clarify how ASR outputs are a function of the decisions made by the training data's human transcribers.

语音识别语言变体转录差异非裔英语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。