arXiv:2601.06307cs.CL2026-01被引 1

用质量评估模型奖励机制,让翻译模型更好处理习语,连带提升整体翻译水平。

A Rising Tide Lifts All Boats: MTQE Rewards for Idioms Improve General Translation Quality

  • 用MTQE模型做奖励信号,指导模型优化习语翻译。
  • 习语翻译得分提升14点,普通文本翻译隐含提升8点。
  • 适合想提升跨文化语言理解能力的研究者参考。

非构词性表达(如习语、谚语、隐喻)因其意义无法从单个词汇推导,对神经机器翻译系统构成重大挑战。这些表达蕴含丰富的文化内涵,兼具字面与比喻意义,翻译难度高。由于模型在构词性文本上表现良好,本文采用类似GRPO的微调策略,以机器翻译质量评估(MTQE)模型作为奖励函数,训练模型更优地翻译习语。基于中文和印地语习语数据集的实验表明,习语翻译能力提升约14分,非习语文本的通用翻译质量隐含提升约8分,跨语言迁移能力(在一种语言上训练,另一种语言上评估)提升约6分。本研究量化了非构词性表达的翻译差距,为发展具备更强跨文化与修辞理解能力的大模型提供了洞见。

原文摘要 · Abstract (English)

Non-compositional expressions (e.g., idioms, proverbs, and metaphors) pose significant challenges for neural machine translation systems because their meanings cannot be derived from individual words alone. These expressions encode rich, cultural meaning, and have both figurative and literal meanings, making accurate translation difficult. Because models are fairly good at translating compositional text, we investigate GRPO-style fine-tuning using Machine Translation Quality Estimation (MTQE) models as reward functions to train models to better translate idioms. Using Chinese and Hindi idiom datasets, we find that idiom translation abilities improve by ~14 points, general, non-idiomatic translation implicitly improves by ~8 points, and cross-lingual translation abilities (trained on one language, evaluated on another) improves by ~6 points. Overall, our work quantifies the non-compositional translation gap and offers insights for developing LLMs with stronger cross-cultural and figurative language understanding.

习语翻译质量评估跨语言强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。