让大模型在翻译中有效推理,需定制化结构化方法
Unlocking Reasoning Capability on Machine Translation in Large Language Models
- 设计分步草稿、校准、润色与选择性修正的结构化推理框架
- 在WMT24++上显著提升翻译质量,优于通用推理注入
- 发现翻译推理缺乏修正与探索,通用推理不适用
推理导向的大语言模型在数学和编程任务中表现优异,但其在机器翻译(MT)中的影响尚未充分探索。我们在WMT24++基准上系统评估多个开源与闭源推理模型,发现开启显式推理反而普遍降低翻译质量,跨语言和模型均如此。分析表明,翻译推理路径高度线性,缺乏修正、自纠错及备选译法探索,实用性有限。即使引入更强模型生成的高质量推理轨迹,也无法稳定提升弱模型性能。为解决这一不匹配问题,我们提出针对翻译优化的结构化推理框架,包含多步草稿、语义完备性精炼、流畅性优化和选择性迭代修正。我们构建了动态结构化推理轨迹的合成数据集,并在此基础上对大型推理模型进行后训练。实验显示,该方法显著优于标准微调和通用推理注入基线。研究证明,推理必须适配任务特性才能真正助力机器翻译。
原文摘要 · Abstract (English)
Reasoning-oriented large language models (RLMs) achieve strong gains on tasks such as mathematics and coding by generating explicit intermediate reasoning. However, their impact on machine translation (MT) remains underexplored. We systematically evaluate several open- and closed-weights RLMs on the WMT24++ benchmark and find that enabling explicit reasoning consistently degrades translation quality across languages and models. Analysis reveals that MT reasoning traces are highly linear, lacking revision, self-correction and exploration of alternative translations, which limits their usefulness. Furthermore, injecting higher-quality reasoning traces from stronger models does not reliably improve weaker models' performance. To address this mismatch, we propose a structured reasoning framework tailored to translation, based on multi-step drafting, adequacy refinement, fluency improvement, and selective iterative revision. We curate a synthetic dataset of dynamic structured reasoning traces and post-train a large reasoning model on this data. Experiments show significant improvements over standard translation fine-tuning and injected generic reasoning baselines. Our findings demonstrate that reasoning must be task-structured to benefit MT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。