探究强化学习中推理过程对机器翻译质量与成本的影响
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
- 训练或推理时移除模型推理痕迹,对比翻译效果
- 推理阶段保留痕迹可提升翻译质量,但增加输出长度
- 揭示了计算开销与翻译质量间的权衡关系
基于可验证奖励的强化学习(RLVR)已成为大语言模型后训练的有效范式,适用于神经机器翻译等下游任务。最新研究表明,由于引入了推理能力,RLVR可能更适合法律文档翻译。然而,这究竟是推理能力带来的优势,还是训练范式本身的作用尚不明确。本文通过系统性地在训练或推理阶段移除模型的推理痕迹,考察其重要性。实验表明,在推理阶段保留推理过程能有效提升整体翻译质量。同时,推理过程会增加输出词元数量,因此我们进一步研究了由此带来的计算开销与翻译质量之间的权衡关系。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest research indicating that RLVR could be the preferred training method for translating legal documents due to the induced reasoning capabilities, it raises the question whether it is really attributed to the reasoning or more generally to the training paradigm. We investigate the importance of including the model's reasoning trace in the generated responses during both training and inference by systematically omitting it from one of the phases. Our experiments show that including the reasoning, specifically during inference, has a positive effect on the overall translation quality. Furthermore, we recognise that the reasoning leads to an increase in output tokens, hence we study the cost-quality tradeoff between the increased computational demands and the improved translation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。