arXiv:2503.04797cs.CLcs.LG2025-03中稿 · NACCL综述被引 13

系统梳理低资源印地语系语言的双语语料库,助力机器翻译发展

Parallel Corpora for Machine Translation in Low-resource Indic Languages: A Comprehensive Review

  • 按文本、混码、多模态分类整理印地语系语料库
  • 揭示语料质量与数据稀疏性对翻译性能的制约
  • 适合研究低资源语言翻译与语料构建的学者参考

双语语料库在训练机器翻译模型中至关重要,尤其对于高质双语数据稀缺的低资源语言。本文全面综述了印地语系语言的可用平行语料库,涵盖多种语言家族、书写系统和区域变体。将语料库分为文本对文本、混码及各类多模态数据,强调其对鲁棒多语言翻译系统发展的意义。除资源统计外,还深入分析语料构建中的挑战:语言多样性、书写差异、数据稀缺以及非正式文本泛滥。评估语料在对齐质量、领域代表性等方面的特性。指出各印地语间数据不平衡、质量与数量权衡、噪声与方言数据对翻译表现的影响等开放问题。最后提出未来方向:利用跨语言迁移学习、扩展多语言数据集、融合多模态资源以提升翻译质量。据我们所知,这是首篇聚焦低资源印地语系语言机器翻译的平行语料库综合性综述。

原文摘要 · Abstract (English)

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel corpora for Indic languages, which span diverse linguistic families, scripts, and regional variations. We categorize these corpora into text-to-text, code-switched, and various categories of multimodal datasets, highlighting their significance in the development of robust multilingual MT systems. Beyond resource enumeration, we critically examine the challenges faced in corpus creation, including linguistic diversity, script variation, data scarcity, and the prevalence of informal textual content.We also discuss and evaluate these corpora in various terms such as alignment quality and domain representativeness. Furthermore, we address open challenges such as data imbalance across Indic languages, the trade-off between quality and quantity, and the impact of noisy, informal, and dialectal data on MT performance. Finally, we outline future directions, including leveraging cross-lingual transfer learning, expanding multilingual datasets, and integrating multimodal resources to enhance translation quality. To the best of our knowledge, this paper presents the first comprehensive review of parallel corpora specifically tailored for low-resource Indic languages in the context of machine translation.

机器翻译低资源语言语料库印地语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。