评估大模型在英泰互译中的表现,发现数据质量与领域匹配至关重要。
Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets
- 对比多模型在多个数据集上的翻译效果,分析不同数据质量的影响。
- 发现数据集质量和领域一致性显著影响模型性能,尤其在低资源语言上。
- 验证提示学习可实现少样本翻译,且生成结果结构合理,适合快速部署。
低资源语言如泰米尔语的机器翻译难题主要源于平行语料有限、领域差异大及形态复杂。本研究系统评估了多种多语言翻译模型在英语-泰米尔语和泰米尔语-英语互译任务上的表现,覆盖NTREX、EnTamV2、WikiMatrix和PMIndia等多个数据集。采用BLEU和chrF指标,比较了监督式NMT模型(NLLB、mBART)在不同质量与领域数据上的表现。通过注意力可视化分析源文与目标文的词元对齐,提升模型可解释性。研究还表明,使用上下文提示(in-context prompting)可借助具备泰米尔语能力的TamilLaMA模型实现少样本英泰/泰英翻译,且生成结果在结构上保持连贯,与监督方法相比具有竞争力。结果说明:数据集质量与领域匹配度直接影响模型表现,注意力机制有助于理解翻译过程,而少样本大模型仍能生成有效翻译。
原文摘要 · Abstract (English)
The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and morphological complexity. This research presents the comprehensive evaluation of the performance of several multilingual translation models on English-Tamil and Tamil-English translations across multiple datasets: NTREX, EnTamV2, WikiMatrix and PMIndia. This study evaluates supervised NMT systems, NLLB and mBART, using both the BLEU and chrF metric, and examines how these systems perform on data of different quality levels and domains. This performs an attention-based analysis to increase model interpretability by visualising the alignments of tokens in an English source text and their Tamil translations and vice-versa to provide insight into how they make translations. This study also demonstrates that using in-context prompting can provide an excellent way to perform a few-shot translation of English to Tamil and Tamil-English using a Tamil capable TamilLaMA model, and compare this to supervised approaches qualitatively. These findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。