用预训练语音模型跨语言识别重音,效果优于传统方法。
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
- 用Transformer语音编码器+微调,替代传统声学特征
- 克罗地亚语和塞尔维亚语识别接近完美,方言下降10分
- 仅需数百个双音节词就能达到良好效果,适合资源少语言
自动化重音识别在语义表达与语音理解中至关重要。以往研究多依赖传统声学特征和英语数据集。本文采用预训练的Transformer模型并添加音频帧分类头进行微调。实验使用新构建的克罗地亚语训练数据集,测试集涵盖克罗地亚语、塞尔维亚语、查卡瓦方言和斯洛文尼亚语。对比基于传统声学特征的SVM分类器,微调后的语音Transformer在所有语言上均表现更优,在克罗地亚语和塞尔维亚语上接近完美,而在较远的查卡瓦方言和斯洛文尼亚语上性能下降约10分。最后,我们发现仅需几百个双音节词即可实现强性能。相关数据集与模型已开源。
原文摘要 · Abstract (English)
Automating primary stress identification has been an active research field due to the role of stress in encoding meaning and aiding speech comprehension. Previous studies relied mainly on traditional acoustic features and English datasets. In this paper, we investigate the approach of fine-tuning a pre-trained transformer model with an audio frame classification head. Our experiments use a new Croatian training dataset, with test sets in Croatian, Serbian, the Chakavian dialect, and Slovenian. By comparing an SVM classifier using traditional acoustic features with the fine-tuned speech transformer, we demonstrate the transformer's superiority across the board, achieving near-perfect results for Croatian and Serbian, with a 10-point performance drop for the more distant Chakavian and Slovenian. Finally, we show that only a few hundred multi-syllabic training words suffice for strong performance. We release our datasets and model under permissive licenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。