arXiv:2601.12199cs.CL2026-01中稿 · IEEE ICASSP 2026被引 1

用语音识别思路做阿拉伯方言识别,适合实时流式应用。

CTC-DID: CTC-Based Arabic dialect identification for streaming applications

  • 将方言识别当作有限词汇的语音识别任务处理。
  • 在小数据上训练的模型优于微调的Whisper和ECAPA-TDNN。
  • 对短语音更鲁棒,适合实时流式部署。

本文提出一种受连接时序分类(CTC)损失启发的阿拉伯方言识别(DID)方法。该方法将方言识别建模为一个有限词汇量的自动语音识别系统,将方言标签视为给定语音片段的标签序列。训练时,通过提出的语言无关启发式(LAH)方法或预训练的ASR模型估计转录中方言标签的重复。在低资源阿拉伯方言识别(ADI)任务上评估,实验表明基于自监督学习(SSL)的CTC-DID模型在有限数据上训练后,性能优于微调的Whisper和ECAPA-TDNN模型。值得注意的是,该方法在Casablanca数据集的零样本测试中也表现更优。此外,该方法对短语音更具鲁棒性,且易于适配流式、实时应用场景,性能衰减极小。

原文摘要 · Abstract (English)

This paper proposes a Dialect Identification (DID) approach inspired by the Connectionist Temporal Classification (CTC) loss function as used in Automatic Speech Recognition (ASR). CTC-DID frames the dialect identification task as a limited-vocabulary ASR system, where dialect tags are treated as a sequence of labels for a given utterance. For training, the repetition of dialect tags in transcriptions is estimated either using a proposed Language-Agnostic Heuristic (LAH) approach or a pre-trained ASR model. The method is evaluated on the low-resource Arabic Dialect Identification (ADI) task, with experimental results demonstrating that an SSL-based CTC-DID model, trained on a limited dataset, outperforms both fine-tuned Whisper and ECAPA-TDNN models. Notably, CTC-DID also surpasses these models in zero-shot evaluation on the Casablanca dataset. The proposed approach is found to be more robust to shorter utterances and is shown to be easily adaptable for streaming, real-time applications, with minimal performance degradation.

方言识别语音识别实时系统自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。