对比大模型与推理模型在多领域翻译中的表现,发现推理模型更擅长复杂长文本翻译。
How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation
- 用结构化推理提升跨领域翻译质量,针对15个领域测试不同翻译方向。
- 在长文本和高难度任务中,推理模型比传统大模型翻译准确率显著更高。
- 通过适配领域的提示词策略,进一步发挥推理模型优势,适合专业翻译场景。
大型语言模型(LLMs)在通用机器翻译中表现优异,但在复杂、领域敏感的翻译任务中效果仍不明确。近期出现的大规模推理模型(LRMs)引发疑问:结构化推理能否提升跨领域翻译质量?本文在15个代表性领域和4种翻译方向上,对比了LRMs与传统LLMs的表现,考察了任务难度、输入长度和术语密度等因素。采用自动评估指标与改进的MQM评估体系进行综合评价。结果表明,LRMs在语义复杂的领域中持续优于传统LLMs,尤其在长文本和高难度场景下表现突出。此外,领域自适应提示策略能进一步释放LRMs的推理潜力。研究揭示了结构化推理在多领域机器翻译中的价值,并为优化领域敏感翻译系统提供了实用洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong performance in general-purpose machine translation, but their effectiveness in complex, domain-sensitive translation tasks remains underexplored. Recent advancements in Large Reasoning Models (LRMs), raise the question of whether structured reasoning can enhance translation quality across diverse domains. In this work, we compare the performance of LRMs with traditional LLMs across 15 representative domains and four translation directions. Our evaluation considers various factors, including task difficulty, input length, and terminology density. We use a combination of automatic metrics and an enhanced MQM-based evaluation hierarchy to assess translation quality. Our findings show that LRMs consistently outperform traditional LLMs in semantically complex domains, especially in long-text and high-difficulty translation scenarios. Moreover, domain-adaptive prompting strategies further improve performance by better leveraging the reasoning capabilities of LRMs. These results highlight the potential of structured reasoning in MDMT tasks and provide valuable insights for optimizing translation systems in domain-sensitive contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。