arXiv:2510.18077cs.CL2025-10被引 2

用思维链提升大模型的上下文翻译能力

Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Models

  • 通过思维链提示让模型分析句子间指代关系
  • 最佳模型在判别任务达90%准确率,生成任务COMET得分92%
  • 能力强的模型经思维链后提升更显著

本文评估大语言模型(LLMs)在处理含句间依赖的文本翻译中的表现。采用English-French DiscEvalMT基准(Bawden et al., 2018),包含涉及代词回指和词汇衔接挑战的句子对。在两个任务上评估来自DeepSeek-R1、GPT、Llama、Mistral和Phi系列的12个LLMs:(1)从正确与看似合理但错误的翻译中识别正确选项;(2)生成正确翻译。对比鼓励思维链推理的提示与不鼓励的提示。表现最佳的模型利用推理后,在第一项任务中达到约90%准确率,第二项任务中COMET得分约92%,其中GPT-4、GPT-4o和Phi表现突出。此外观察到“越强越强”效应:原本表现好的模型通过推理获得的提升更大。

原文摘要 · Abstract (English)

This paper assesses the ability of large language models (LLMs) to translate texts that include inter-sentential dependencies. We use the English-French DiscEvalMT benchmark (Bawden et al., 2018) with pairs of sentences containing translation challenges for pronominal anaphora and lexical cohesion. We evaluate 12 LLMs from the DeepSeek-R1, GPT, Llama, Mistral and Phi families on two tasks: (1) distinguish a correct translation from a wrong but plausible one; and (2) generate a correct translation. We compare prompts that encourage chain-of-thought reasoning with those that do not. The best models take advantage of reasoning and reach about 90% accuracy on the first task and COMET scores of about 92% on the second task, with GPT-4, GPT-4o and Phi standing out. Moreover, we observe a "wise get wiser" effect: the improvements through reasoning are larger for models that already perform well without reasoning.

机器翻译思维链大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。