arXiv:2606.18597cs.CL2026-06被引 8

用迁移学习和数据增强提升低资源方言识别性能

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

  • 基于源方言预训练模型,通过音高、速度等增强目标方言数据
  • 在两个基准数据集上显著超越现有方法,准确率提升明显
  • 适合资源匮乏的方言识别研究者参考

由于标注资源稀缺,中文方言识别是一项具有挑战性的自然语言处理任务。本文提出一种结合迁移学习与数据增强的中文方言识别框架(CDDTLDA),以应对资源不足问题。首先,利用较大规模的中文方言语料库训练一个源端自动语音识别(ASR)模型;随后,采用速度、音高和噪声扰动等简单但有效的数据增强方法,扩充目标端低资源方言数据,并基于源端模型微调另一个目标端ASR模型。同时,通过自注意力机制捕捉源端与目标端模型间的潜在共性语义特征。最后,提取目标ASR模型中的隐含语义表示,用于中文方言识别。大量实验表明,该模型在两个基准中文方言语料库上显著优于当前最优方法。

原文摘要 · Abstract (English)

Chinese dialects discrimination is a challenging natural language processing task due to scarce annotation resource. In this article, we develop a novel Chinese dialects discrimination framework with transfer learning and data augmentation (CDDTLDA) in order to overcome the shortage of resources. To be more specific, we first use a relatively larger Chinese dialects corpus to train a source-side automatic speech recognition (ASR) model. Then, we adopt a simple but effective data augmentation method (i.e., speed, pitch, and noise disturbance) to augment the target-side low-resource Chinese dialects, and fine-tune another target ASR model based on the previous source-side ASR model. Meanwhile, the potential common semantic features between source-side and target-side ASR models can be captured by using self-attention mechanism. Finally, we extract the hidden semantic representation in the target ASR model to conduct Chinese dialects discrimination. Our extensive experimental results demonstrate that our model significantly outperforms state-of-the-art methods on two benchmark Chinese dialects corpora.

方言识别迁移学习数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。