让AI翻译时‘重译一遍’比步步分解更有效,挑战了人类思维的翻译范式。
Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation
- 让大模型重译并自我修正,比刻意分步推理效果更好
- 在WMT24测试集上,重译策略超越当前最优的分步提示方法
- 揭示大模型与人类在翻译最优策略上的根本差异
大型语言模型(LLMs)在诸多任务中展现出强大的推理能力,常通过链式思维(CoT)明确分解任务。近期基于LLM的翻译研究设计了人工构造的分步提示,或训练模型加入中间步骤。例如,Translating Step-by-step(Briakou et al., 2024)提出多步提示,包含分解与精炼过程,在WMT24测试数据上达到领先性能。本文对此策略的有效性进行审视:实证发现,至少在测试模型上,并无明确证据表明性能提升源于显式的分步推理;反而,引导模型‘重译’并自我修正,其效果优于人类式分步提示。尽管分解影响翻译行为,但对分解的忠实度既有正面也有负面作用。分析表明,人类与大模型在最优翻译策略上存在分歧。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate strong reasoning capabilities for many tasks, often by explicitly decomposing the task via Chain-of-Thought (CoT) reasoning. Recent work on LLM-based translation designs hand-crafted prompts to decompose translation, or trains models to incorporate intermediate steps. Translating Step-by-step (Briakou et al., 2024), for instance, introduces a multi-step prompt with decomposition and refinement of translation with LLMs, which achieved state-of-the-art results on WMT24 test data. In this work, we scrutinise this strategy's effectiveness. Empirically, we find no clear evidence that performance gains stem from explicitly decomposing the translation process via CoT, at least for the models on test; and we show prompting LLMs to 'translate again' and self-refine yields even better results than human-like step-by-step prompting. While the decomposition influences translation behaviour, faithfulness to the decomposition has both positive and negative effects on translation. Our analysis therefore suggests a divergence between the optimal translation strategies for humans and LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。