arXiv:2605.24451cs.CL2026-05

为越南语方言发音差异建模,提升语音识别准确率。

Phonetic Modeling of Dialectal Variation in Vietnamese Speech

论文配图:Phonetic Modeling of Dialectal Variation in Vietnamese Speech
图 1 · 摘自论文原文
  • 构建音素级词汇表,分解音节为结构化发音成分
  • 在多方言数据集上性能媲美最强预训练模型,参数更少
  • 适合需要低资源方言语音识别的开发者与研究者

越南语在北方、中部和南方存在显著的方言发音差异,相同词项可能有明显不同的发音表现。这种差异给自动语音识别(ASR)带来挑战,且由于越南语拼写与发音关系复杂,计算建模困难。现有方法通常在词级别处理方言差异,假设拼写与发音映射不变,难以捕捉系统性发音差异。本文提出一种方言感知的发音建模框架,显式建模越南语音系结构与方言变异,涵盖词汇与解码两个层面。该框架引入音素级词汇表,将每个音节分解为结构化发音成分,并映射到特定方言的国际音标(IPA)表示;同时设计音素结构解码器,联合预测这些成分。在唯一可用的多方言越南语数据集UIT-ViMD上的实验表明,该方法优于多种预训练基线模型,尤其在跨方言测试中性能与最强预训练wav2vec2-base-vi-250h相当,但使用参数更少且无需外部预训练。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different pronunciations. Such variation poses challenges for automatic speech recognition (ASR) and remains difficult to model computationally due to the complex relationship between Vietnamese orthography and phonology. Existing approaches typically address dialect variability at the word level, assuming dialect-invariant mappings between spelling and pronunciation, which limits their ability to capture systematic phonetic differences. We propose a dialect-aware phonetic framework that explicitly models Vietnamese phonological structure and dialectal variation at both the vocabulary and decoding levels. The framework introduces a phonetic vocabulary that decomposes each syllable into structured phonetic components and maps them to dialect-specific IPA representations, together with a phonetic-structure decoder that jointly predicts these components. Experiments on the UIT-ViMD, a only-available dataset for multi-dialect in Vietnamese, show that the proposed approach outperforms various pre-trained baselines, \textbf{especially matches the performance of the strongest pretrained wav2ve2-base-vi-250h} across dialects while \textbf{using substantially fewer parameters and no external pretraining}. Code for experimental reproducibility will be publicly available upon the acceptance of this paper.

语音识别方言建模音素结构低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。