将翻译拆解为多步流程,提升长文本翻译质量。
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
- 把翻译分为调研、初稿、润色、校对四步,逐步优化。
- 在10个语种对上超越传统零样本提示,达WMT2024最佳水平。
- 适合需要高质量长文本翻译的研究者与从业者。
本文提出一种分步式长文本翻译方法,借鉴翻译研究中的成熟流程。不同于将机器翻译视为单一任务,我们设计了一个多轮交互框架,让语言模型依次完成译前调研、初稿撰写、内容润色和校对,实现翻译质量的逐步提升。在十个语种对上,使用Gemini 1.5 Pro进行大量自动评估,结果表明分步翻译显著优于传统的零样本提示方法及早期类人基线策略,在WMT2024上达到当前最优表现。
原文摘要 · Abstract (English)
In this paper we present a step-by-step approach to long-form text translation, drawing on established processes in translation studies. Instead of viewing machine translation as a single, monolithic task, we propose a framework that engages language models in a multi-turn interaction, encompassing pre-translation research, drafting, refining, and proofreading, resulting in progressively improved translations. Extensive automatic evaluations using Gemini 1.5 Pro across ten language pairs show that translating step-by-step yields large translation quality improvements over conventional zero-shot prompting approaches and earlier human-like baseline strategies, resulting in state-of-the-art results on WMT2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。