用细粒度错误严重度优化机器翻译,提升质量与训练稳定性
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
- 以词元级错误严重度为奖励信号,替代传统句子级反馈
- 在多语言数据集上,自动与人工评估均优于基线模型
- 训练过程更稳定,奖励值随训练轮次持续上升
强化学习(RL)已被证明是训练神经机器翻译系统的一种有效且稳健的方法,尤其当搭配能准确评估翻译质量的强大奖励模型时。然而,大多数研究集中于使用句子级反馈的强化学习方法,导致学习信号效率低下,因奖励稀疏性问题——整个句子仅获得单一评分。为此,我们提出一种新方法,利用细粒度的词元级质量评估及错误严重度等级,结合强化学习技术。具体而言,我们采用xCOMET这一最先进的质量评估系统作为词元级奖励模型。我们在小型和大型翻译数据集上,对基于标准编码器-解码器和大语言模型的机器翻译系统进行实验,比较句子级与细粒度奖励信号对翻译质量的影响。结果表明,使用词元级奖励训练可显著提升多种语言对的翻译质量,无论是自动评估还是人工评估均优于基线。此外,词元级奖励优化使训练更加稳定,表现为平均奖励值在训练周期中稳步增长。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has been proven to be an effective and robust method for training neural machine translation systems, especially when paired with powerful reward models that accurately assess translation quality. However, most research has focused on RL methods that use sentence-level feedback, leading to inefficient learning signals due to the reward sparsity problem -- the model receives a single score for the entire sentence. To address this, we propose a novel approach that leverages fine-grained, token-level quality assessments along with error severity levels using RL methods. Specifically, we use xCOMET, a state-of-the-art quality estimation system, as our token-level reward model. We conduct experiments on small and large translation datasets with standard encoder-decoder and large language models-based machine translation systems, comparing the impact of sentence-level versus fine-grained reward signals on translation quality. Our results show that training with token-level rewards improves translation quality across language pairs over baselines according to both automatic and human evaluation. Furthermore, token-level reward optimization improves training stability, evidenced by a steady increase in mean rewards over training epochs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。