通过筛选源语言复杂句,用更少数据提升低资源翻译质量
Get away with less: Need of source side data curation to build parallel corpus for low resource Machine Translation
- 基于词汇与语言特征筛选源句,构建高效平行语料
- 50K至800K句规模下均显著提升翻译效果
- 可减少超一半训练数据,适合低资源语言场景
数据清洗是机器翻译训练中的关键但研究不足环节。低资源语言缺乏足够人工翻译数据,导致训练成本过高。本文提出LALITA框架,通过词汇和语言学特征筛选源语言句子,构建高效平行语料。实验以英-印地语双语语料为基础,验证在50K至800K英文句规模下,仅使用复杂句即可显著提升翻译质量。该方法在印地语、奥里亚语、尼泊尔语、挪威新诺尔思克语及德语中均实现训练数据需求减少超50%,有效降低训练成本,并具备数据增强潜力。
原文摘要 · Abstract (English)
Data curation is a critical yet under-researched step in the machine translation training paradigm. To train translation systems, data acquisition relies primarily on human translations and digital parallel sources or, to a limited degree, synthetic generation. But, for low-resource languages, human translation to generate sufficient data is prohibitively expensive. Therefore, it is crucial to develop a framework that screens source sentences to form efficient parallel text, ensuring optimal MT system performance in low-resource environments. We approach this by evaluating English-Hindi bi-text to determine effective sentence selection strategies for optimal MT system training. Our extensively tested framework, (Lexical And Linguistically Informed Text Analysis) LALITA, targets source sentence selection using lexical and linguistic features to curate parallel corpora. We find that by training mostly on complex sentences from both existing and synthetic datasets, our method significantly improves translation quality. We test this by simulating low-resource data availabilty with curated datasets of 50K to 800K English sentences and report improved performances on all data sizes. LALITA demonstrates remarkable efficiency, reducing data needs by more than half across multiple languages (Hindi, Odia, Nepali, Norwegian Nynorsk, and German). This approach not only reduces MT systems training cost by reducing training data requirement, but also showcases LALITA's utility in data augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。