arXiv:2510.06471cs.CL2025-10Conference of the …被引 4

通过推理模型的测试时扩展,提升机器翻译质量。

Test-Time Scaling of Reasoning Models for Machine Translation

  • 在推理阶段增加计算量,探索翻译性能提升。
  • 领域微调后,推理深度可带来持续改进,超限则性能下降。
  • 后编辑场景中,自纠错机制显著受益于测试时扩展。

测试时扩展(TTS)已提升各类任务中推理模型(RMs)的表现,如数学和编程,但在机器翻译(MT)中的效果尚未充分研究。本文评估了12个推理模型在多个领域的多样化机器翻译基准上的表现,涵盖直接翻译、强制推理外推和后编辑三种场景。结果表明,通用型模型在直接翻译中,TTS带来的收益有限且不稳定,性能很快达到平台期。而经过领域特定微调后,模型的推理过程与任务需求对齐,使推理深度可实现一致提升,直至自适应最优值。强制模型超出自然停止点推理会持续降低翻译质量。相反,在后编辑场景中,TTS能有效将自我修正转化为有益过程。研究揭示:在机器翻译中,推理时计算的价值不在于提升通用模型的单次翻译,而在于多步自纠正工作流及与任务专用模型结合的应用。

原文摘要 · Abstract (English)

Test-time scaling (TTS) has enhanced the performance of Reasoning Models (RMs) on various tasks such as math and coding, yet its efficacy in machine translation (MT) remains underexplored. This paper investigates whether increased inference-time computation improves translation quality. We evaluate 12 RMs across a diverse suite of MT benchmarks spanning multiple domains, examining three scenarios: direct translation, forced-reasoning extrapolation, and post-editing. Our findings show that for general-purpose RMs, TTS provides limited and inconsistent benefits for direct translation, with performance quickly plateauing. However, the effectiveness of TTS is unlocked by domain-specific fine-tuning, which aligns a model's reasoning process with task requirements, leading to consistent improvements up to an optimal, self-determined reasoning depth. We also find that forcing a model to reason beyond its natural stopping point consistently degrades translation quality. In contrast, TTS proves highly effective in a post-editing context, reliably turning self-correction into a beneficial process. These results indicate that the value of inference-time computation in MT lies not in enhancing single-pass translation with general models, but in targeted applications like multi-step, self-correction workflows and in conjunction with task-specialized models.

机器翻译推理模型测试时扩展后编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。