arXiv:2510.13211cs.CL2025-10

用图文分析自动构建低资源语言平行语料库,提升机器翻译效果。

A fully automated and scalable Parallel Data Augmentation for Low Resource Languages using Image and Text Analytics

  • 通过图像与文本分析从报纸中自动提取双语语料
  • 在两种语言组合上构建语料库,机器翻译提升近3 BLEU点
  • 全流程自动化且可扩展,适合数据稀缺语言研究

全球语言多样性导致数字语言资源分布不均,限制了多数人类群体的技术受益。低资源语言缺乏数据支持,难以开展自然语言处理任务。本文提出一种新颖的可扩展、全自动方法,利用图像与文本分析技术,从报纸文章中提取双语平行语料库。我们在两种不同语言组合上验证该方法,构建平行语料库,并通过下游机器翻译任务证明其有效性,相较当前基线提升接近3 BLEU点。

原文摘要 · Abstract (English)

Linguistic diversity across the world creates a disparity with the availability of good quality digital language resources thereby restricting the technological benefits to majority of human population. The lack or absence of data resources makes it difficult to perform NLP tasks for low-resource languages. This paper presents a novel scalable and fully automated methodology to extract bilingual parallel corpora from newspaper articles using image and text analytics. We validate our approach by building parallel data corpus for two different language combinations and demonstrate the value of this dataset through a downstream task of machine translation and improve over the current baseline by close to 3 BLEU points.

数据增强低资源语言机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。