arXiv:2504.05914cs.CL2025-04

用合成数据提升英文到泰卢固语翻译,让资源匮乏语言也能用上高质量翻译。

High-Resource Translation:Turning Abundance into Accessibility

  • 用回译生成合成语料,扩充低资源语言数据集
  • 结合预训练模型与参数优化,显著提升翻译质量
  • 适合关注小语种翻译落地的开发者与研究者

本文提出一种构建英文至泰卢固语翻译模型的新方法,利用迁移学习应对低资源语言挑战。以Bharat Parallel Corpus Collection(BPCC)为主要数据集,通过迭代回译生成合成平行语料,有效扩充训练数据并增强模型能力。研究聚焦于数据增强、训练参数优化及预训练模型的高效利用,形成一套综合策略,旨在提升模型对复杂句式与语言细微差别的处理能力。该工作强调创新数据处理技术与迁移学习在克服稀疏数据限制中的价值,为机器翻译领域提供新思路,助力英-泰卢固语实际交流场景中的沟通效率提升。

原文摘要 · Abstract (English)

This paper presents a novel approach to constructing an English-to-Telugu translation model by leveraging transfer learning techniques and addressing the challenges associated with low-resource languages. Utilizing the Bharat Parallel Corpus Collection (BPCC) as the primary dataset, the model incorporates iterative backtranslation to generate synthetic parallel data, effectively augmenting the training dataset and enhancing the model's translation capabilities. The research focuses on a comprehensive strategy for improving model performance through data augmentation, optimization of training parameters, and the effective use of pre-trained models. These methodologies aim to create a robust translation system that can handle diverse sentence structures and linguistic nuances in both English and Telugu. This work highlights the significance of innovative data handling techniques and the potential of transfer learning in overcoming limitations posed by sparse datasets in low-resource languages. The study contributes to the field of machine translation and seeks to improve communication between English and Telugu speakers in practical contexts.

机器翻译低资源语言数据增强回译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。