用wav2vec2.0自动识别辅音爆破音,准确率高且适用多种语料。
Automatic classification of stop realisation with wav2vec2.0
- 基于wav2vec2.0模型,训练分类爆破音是否存在。
- 在英日语数据中准确率达高,跨语种和语料质量均稳定。
- 结果接近人工标注,适合语音研究规模化分析。
现代语音研究常依赖自动化工具进行语音数据标注,但针对诸多可变语音现象的标注工具仍稀缺。预训练自监督模型如wav2vec2.0在语音分类任务中表现优异,并能隐式编码细粒度语音信息。本文展示wav2vec2.0可在英语和日语中高精度自动分类爆破音的存在与否,且在精细标注与未准备语料上均表现稳健。自动标注结果复现了爆破音实现的变异模式,与人工标注高度一致。这表明预训练语音模型具备成为语音语料自动化标注与处理工具的巨大潜力,使语音研究可轻松实现规模扩展。
原文摘要 · Abstract (English)
Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic information. We demonstrate that wav2vec2.0 models can be trained to automatically classify stop burst presence with high accuracy in both English and Japanese, robust across both finely-curated and unprepared speech corpora. Patterns of variability in stop realisation are replicated with the automatic annotations, and closely follow those of manual annotations. These results demonstrate the potential of pre-trained speech models as tools for the automatic annotation and processing of speech corpus data, enabling researchers to 'scale-up' the scope of phonetic research with relative ease.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。