arXiv:2502.19735cs.CL2025-02Transactions of th…被引 20

用强化学习让大模型自动生成人类级翻译推理链,提升跨语言翻译能力。

R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning

  • 通过强化学习驱动模型自动生成符合人类习惯的六类推理模板。
  • 在10+语言、40+方向上实现稳定性能提升,未见语言表现更优。
  • 适合需要高适应性翻译的场景,如多语种、跨领域任务。

尽管近期推理增强型大语言模型(如DeepSeek-R1)取得进展,但在机器翻译中引入推理时推理链(CoT)仍属空白。现有方法或针对特定子任务设计固定推理链,或依赖与人类不一致的合成数据及易过拟合的监督微调,限制了对多样翻译场景的适应性。本文提出R1-Translator(R1-T1),一种基于强化学习的推理框架,实现通用机器翻译中的推理时推理生成。该框架首次实现三项创新:(1) 将推理式翻译扩展至训练阶段未见的广泛场景(如多语言、领域特定翻译);(2) 明确定义六类专家标注的推理模板,对应人类混合策略如上下文感知改写和反向翻译;(3) 通过强化学习实现推理链的自我演化发现。人工与自动评估均显示,在Flores-101测试集和四个领域特定翻译任务上,覆盖10+语言、40+翻译方向,性能持续提升,尤其在训练中未见的语言上表现显著增强。

原文摘要 · Abstract (English)

Despite recent breakthroughs in reasoning-enhanced large language models (LLMs) like DeepSeek-R1, incorporating inference-time reasoning into machine translation (MT), where human translators naturally employ structured, multi-layered reasoning chain-of-thoughts (CoTs), is yet underexplored. Existing methods either design a fixed CoT tailored for a specific MT sub-task (e.g., literature translation), or rely on synthesizing CoTs unaligned with humans and supervised fine-tuning (SFT) prone to overfitting, limiting their adaptability to diverse translation scenarios. This paper introduces R1-Translator (R1-T1), a novel framework to achieve inference-time reasoning for general MT via reinforcement learning (RL) with human-aligned CoTs comprising six common patterns. Our approach pioneers three innovations: (1) extending reasoning-based translation to broader MT scenarios (e.g., multilingual MT, domain MT) unseen in the training phase; (2) formalizing six expert-curated CoT templates that mirror hybrid human strategies like context-aware paraphrasing and back translation; and (3) enabling self-evolving CoT discovery through RL. Both human and automatic evaluation results indicate a steady translation performance improvement in a total of 10+ languages and 40+ translation directions on Flores-101 test set and four domain-specific MT tasks, especially on the languages unseen from training.

机器翻译推理增强强化学习多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。