arXiv:2510.18898cs.CLcs.CY2025-10被引 2

用Transformer改进孟加拉语到锡尔赫特语的低资源翻译

Transformer-Based Low-Resource Language Translation: A Study on Standard Bengali to Sylheti

  • 微调多语言Transformer模型提升翻译质量
  • mBART-50在语义准确度上最优,MarianMT字符级保真最强
  • 适合关注小语种翻译与包容性技术的研究者

机器翻译已从基于规则和统计的方法演进至以Transformer架构为基础的神经方法。尽管这些方法在高资源语言上取得显著成果,但像锡尔赫特语这样的低资源方言仍研究不足。本文通过微调多语言Transformer模型,并与零样本大语言模型(LLMs)进行比较,研究孟加拉语到锡尔赫特语的翻译。实验表明,微调模型显著优于LLMs:mBART-50在翻译适切性上表现最佳,MarianMT在字符级保真度上最强。结果凸显了针对未充分代表语言进行任务特定适配的重要性,助力包容性语言技术的发展。

原文摘要 · Abstract (English)

Machine Translation (MT) has advanced from rule-based and statistical methods to neural approaches based on the Transformer architecture. While these methods have achieved impressive results for high-resource languages, low-resource varieties such as Sylheti remain underexplored. In this work, we investigate Bengali-to-Sylheti translation by fine-tuning multilingual Transformer models and comparing them with zero-shot large language models (LLMs). Experimental results demonstrate that fine-tuned models significantly outperform LLMs, with mBART-50 achieving the highest translation adequacy and MarianMT showing the strongest character-level fidelity. These findings highlight the importance of task-specific adaptation for underrepresented languages and contribute to ongoing efforts toward inclusive language technologies.

机器翻译低资源Transformer小语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。