构建首个越英混用语平行语料库,提升低资源混用翻译性能
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
- 创建自然出现的越英混用平行语料库VietMix,专家标注
- 数据增强后模型比回译基线高3.5分(xCOMET),零样本提升11.9分
- 框架可迁移,适用于其他低资源混用语言对
机器翻译系统在面对混用语时普遍性能下降,尤其对缺乏专用平行语料的低资源语言更为严重。本文针对越英混用语这一挑战性场景,提出首个由专家翻译、自然产生的越英混用平行语料库VietMix。该语料库解决了正字法模糊和非正式文本中声调符号常被省略的问题。通过构建基于迭代微调与定向筛选的数据增强流水线,验证了VietMix的有效性。实验表明,使用该数据增强的模型相比强回译基线最高提升3.5 xCOMET分,零样本模型性能提升达11.9分。本工作为越英混用翻译提供基础资源,并给出可复用的语料构建与增强框架,适用于其他低资源混用语言场景。
原文摘要 · Abstract (English)
Machine translation (MT) systems universally degrade when faced with code-mixed text. This problem is more acute for low-resource languages that lack dedicated parallel corpora. This work directly addresses this gap for Vietnamese-English, a language context characterized by challenges including orthographic ambiguity and the frequent omission of diacritics in informal text. We introduce VietMix, the first expert-translated, naturally occurring parallel corpus of Vietnamese-English code-mixed text. We establish VietMix's utility by developing a data augmentation pipeline that leverages iterative fine-tuning and targeted filtering. Experiments show that models augmented with our data outperform strong back-translation baselines by up to +3.5 xCOMET points and improve zero-shot models by up to +11.9 points. Our work delivers a foundational resource for a challenging language pair and provides a validated, transferable framework for building and augmenting corpora in other low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。