用少量人工标注+ASR预训练,提升语音转录对非标准口音的鲁棒性
Scaling Human and G2P Supervision for Robust Phonetic Transcription
- 当人工标注少于20-30小时时,自动拼音标注能有效辅助
- 超过此阈值后,自动标注反而降低跨方言泛化能力
- 采用ASR预训练可使错误率降低2.3倍,尤其改善非母语和失语者语音
专家级语音标注成本高昂,尤其针对非标准方言和异常发音。常用替代方案是使用图符到音素(G2P)模型从文本自动生成音素标签以实现大规模标注。本文研究了在英语语音中,人工与G2P标注规模对自动语音转录性能的影响。基于一个涵盖母语、非母语及中风后语言障碍者的80小时精选基准数据集,我们发现存在一个标注质量阈值:当人工标注不足20-30小时时,G2P标注能带来帮助;超过该阈值后,其效果不再显著,甚至降低跨方言鲁棒性。在此之后,使用ASR预训练可实现比以往系统低2.3倍的加权音素特征错误率,并在非母语和失语者语音上取得显著提升。结果表明,单纯依赖数量驱动的G2P标注可能在鲁棒泛化上带来边际收益递减。
原文摘要 · Abstract (English)
Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech. A common alternative is using Grapheme-to-Phoneme (G2P) models to auto-generate phonetic labels from text transcripts at scale. We study how automatic phonetic transcription performance scales with human and G2P supervision in English. Using a curated 80-hour benchmark spanning native, non-native and post-stroke speech, we identify a supervision quality threshold: G2P supervision helps only when fewer than 20-30 hours of human annotation are available. Beyond this threshold, it provides no significant benefit and can reduce cross-dialect robustness. What is effective after this threshold is ASR pretraining which we use to achieve a 2.3x reduction in weighted phone feature error rate over prior systems, with strong gains on non-native and aphasic speech. These results suggest that quantity-driven G2P scaling may yield diminishing returns for robust generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。