针对中文到小语种翻译数据少、质量差的问题,提出新训练框架提升翻译效果。
MERIT: Multilingual Expert-Reward Informed Tuning for Chinese-Centric Low-Resource Machine Translation
- 用特定标记前缀+奖励引导优化,提升小语种译中文性能
- 在5种东南亚低资源语言上显著优于传统方法,超越纯模型放大
- 适合研究中文跨语言翻译或资源匮乏场景的开发者
从中文到低资源东南亚语言的神经机器翻译仍受制于干净双语语料极度稀缺和现有挖掘数据噪声严重。这种长期短缺不仅阻碍模型有效训练,还导致与高资源方向存在巨大性能差距,使老挝语、缅甸语、他加禄语等数百万使用者持续面临低质量翻译系统。本文提出多语言专家-奖励引导微调框架MERIT,将传统的以英语为中心的ALT基准转换为以中文为中心的评估体系,覆盖五种东南亚低资源语言(LRLs)。该框架结合语言特定标记前缀(LTP)、监督微调(SFT)与新型组相对策略优化(GRPO),由语义对齐奖励(SAR)指导。结果表明,在LRL→中文翻译中,针对性数据清洗与奖励引导优化,远胜于单纯模型扩展。
原文摘要 · Abstract (English)
Neural machine translation (NMT) from Chinese to low-resource Southeast Asian languages remains severely constrained by the extreme scarcity of clean parallel corpora and the pervasive noise in existing mined data. This chronic shortage not only impedes effective model training but also sustains a large performance gap with high-resource directions, leaving millions of speakers of languages such as Lao, Burmese, and Tagalog with persistently low-quality translation systems despite recent advances in large multilingual models. We introduce \textbf{M}ultilingual \textbf{E}xpert-\textbf{R}eward \textbf{I}nformed \textbf{T}uning (\textbf{MERIT}), a unified translation framework that transforms the traditional English-centric ALT benchmark into a Chinese-centric evaluation suite for five Southeast Asian low-resource languages (LRLs). Our framework combines language-specific token prefixing (LTP) with supervised fine-tuning (SFT) and a novel group relative policy optimization (GRPO) guided by the semantic alignment reward (SAR). These results confirm that, in LRL{\textrightarrow}Chinese translation, targeted data curation and reward-guided optimization dramatically outperform mere model scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。