arXiv:2603.10211cs.CL2026-03AAAI被引 2

首个覆盖越南全境方言的语料库,助力低资源方言转标准语。

ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation

  • 构建涵盖63个省份的真实语料,包含南北中三方言。
  • 标注一致率超80%,模型在该数据集上最高得分为BLEU 0.8166。
  • 适合做越南语方言处理、低资源NLP研究者使用。

越南方言差异显著,主流NLP系统多基于标准语训练,对非标准方言表现不佳,尤其在中部和南部地区。此前的研究仅聚焦于中部向北部的方言转换,依赖合成数据且多样性不足,忽略了南部方言及北部内部变体。本文提出ViDia2Std,首个覆盖全部63个省份的方言转标准语平行语料库,包含来自真实社交平台评论的超过13,000句对,由三大方言区母语者手工标注。为评估标注一致性,定义了语义映射一致率指标,结果显示北、中、南地区一致性分别为86%、82%、85%。在该语料库上,mBART-large-50表现最佳(BLEU 0.8166,ROUGE-L 0.9384,METEOR 0.8925),ViT5-base则以更少参数实现良好效果。实验表明,方言归一化可显著提升下游任务性能,凸显构建方言感知型资源的重要性。

原文摘要 · Abstract (English)

Vietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predominantly on standard Vietnamese. Such systems often underperform on dialectal inputs, especially from underrepresented Central and Southern regions. Previous work on dialect normalization has focused narrowly on Central-to-Northern dialect transfer using synthetic data and limited dialectal diversity. These efforts exclude Southern varieties and intra-regional variants within the North. We introduce ViDia2Std, the first manually annotated parallel corpus for dialect-to-standard Vietnamese translation covering all 63 provinces. Unlike prior datasets, ViDia2Std includes diverse dialects from Central, Southern, and non-standard Northern regions often absent from existing resources, making it the most dialectally inclusive corpus to date. The dataset consists of over 13,000 sentence pairs sourced from real-world Facebook comments and annotated by native speakers across all three dialect regions. To assess annotation consistency, we define a semantic mapping agreement metric that accounts for synonymous standard mappings across annotators. Based on this criterion, we report agreement rates of 86% (North), 82% (Central), and 85% (South). We benchmark several sequence-to-sequence models on ViDia2Std. mBART-large-50 achieves the best results (BLEU 0.8166, ROUGE-L 0.9384, METEOR 0.8925), while ViT5-base offers competitive performance with fewer parameters. ViDia2Std demonstrates that dialect normalization substantially improves downstream tasks, highlighting the need for dialect-aware resources in building robust Vietnamese NLP systems.

方言识别语料库NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。