用语音特征提升中文方言细粒度识别准确率。
Speech-Driven End-to-End Language Discrimination towards Chinese Dialects
- 融合语音MFCC与词级嵌入的端到端模型
- 在两个方言语料库上优于现有方法
- 适合需要区分相近方言的研究者
相似语言、方言间的语言判别是自然语言处理中的难题,传统文本驱动方法效果不佳。本文探索语音驱动特征在中文方言判别中的有效性。首先系统评估基于CNN的语音MFCC特征适用性;其次设计基于HMM-DNN的端到端语音识别模型,利用注意力机制提取与方言相关的关键词;最后通过卷积神经网络融合词级嵌入与MFCC特征。在两个基准中文方言语料库上的实验表明,所提语音驱动方法在细粒度方言判别上显著优于当前最优方法。
原文摘要 · Abstract (English)
Language discrimination among similar languages, varieties, and dialects is a challenging natural language processing task. The traditional text-driven focus leads to poor results. In this paper, we explore the effectiveness of speech-driven features towards language discrimination among Chinese dialects. First, we systematically explore the appropriateness of speech-driven MFCC features towards CNN-based language discrimination. Then, we design an end-to-end speech recognition model based on HMM-DNN to predict Chinese dialect words. We adopt attention to extract the discriminative words related to different Chinese dialects. Finally, through a CNN, we combine the word-level embedding and the MFCC-based features. Evaluation of two benchmark Chinese dialect corpora shows the appropriateness and effectiveness of the proposed speech-driven approach to fine-grained Chinese dialect discrimination compared to the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。