研究意大利方言语音识别性能差异,发现越接近标准语的方言表现越好。
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
- 用地理与语言相似性分析方言内部差异
- 发现方言识别准确率与标准语相似度相关系数达-0.5
- 可利用地理位置预测未见地点的零样本性能
自然语言处理中对方言的兴趣日益增长,但多数研究仍将方言视为离散类别。例如,英语变体研究常将印度英语或非裔美国黑人口语英语当作同质整体(Faisal et al., 2024;Ziems et al., 2023),然而同一变体内部仍存在显著差异。本文研究意大利方言内部差异,测量其语音识别性能,并实证观察到明显的地理性能差异。该差异与最高性能方言的语言相似度呈显著负相关(相关系数-0.5)。通过对比方言计量方法,发现性能差异源于模型对更接近标准语的方言存在偏倚。此外,借助地统计学方法预测未见地点的零样本性能,结果表明引入地理信息可显著提升预测效果,说明性能分布具有地理结构。
原文摘要 · Abstract (English)
There is increasing interest in looking at dialects in NLP. However, most work to date still treats dialects as discrete categories. For instance, evaluative work in variation-oriented NLP for English often works with Indian English or African-American Venacular English as homogeneous categories (Faisal et al., 2024; Ziems et al., 2023), yet even within one variety there is substantial variation. We examine within-dialect variation and show that performance critically varies within categories. We measure speech-to-text performance on Italian dialects, and empirically observe a geographical performance disparity. This disparity correlates substantially (-0.5) with linguistic similarity to the highest performing dialect variety. We cross-examine our results against dialectometry methods, and interpret the performance disparity to be due to a bias towards dialects that are more similar to the standard variety in the speech-to-text model examined. We additionally leverage geostatistical methods to predict zero-shot performance at unseen sites, and find the incorporation of geographical information to substantially improve prediction performance, indicating there to be geographical structure in the performance distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。