让翻译模型像人一样按难易程度调整思考深度,提升质量和效率。
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

- 根据文本难度自动切换快速直觉与深度推理模式。
- 在15个基准上表现超越更大模型,令牌使用量减少32%至60%。
- 适合需要高效高质多领域翻译的场景,如跨语言客服、内容本地化。
多领域机器翻译(MDMT)因领域间语言复杂度差异而面临挑战。受人类译者根据难度调节思考力度的启发,我们提出TwT(Translation with Thought),一种资源理性框架,可学习在直觉与深思熟虑推理间动态调节。训练分两阶段:(1) 在由DeepSeek-R1生成并经GPT-4o重写以体现人类推理经济性的难度感知长链思维轨迹上进行监督微调;(2) 采用混合奖励的强化学习优化翻译质量与推理效率。在覆盖域内与域外设置、3种已见语言和59种未见语言的15个基准上评估,结果表明:在三种主干模型上,TwT-7B与TwT-14B在翻译质量上优于更大规模的SOTA推理模型,同时令牌使用量降低32%–60%。结果验证了将翻译行为与认知原则对齐,可实现稳健泛化、高质量翻译与高效推理。
原文摘要 · Abstract (English)
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-thought traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32--60\%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。