arXiv:2601.19410cs.CLcs.HC2026-01

大模型虽能自动纠错,但未必用好上下文。

Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?

  • 对比闭源与开源大模型在文档级提示下的纠错能力
  • 闭源模型接近人工水平,但不依赖文档上下文
  • 自动评估指标不可靠,需人工评测

自动后编辑(APE)旨在通过修正机器翻译残余错误来提升译文质量。尽管近期大语言模型(LLMs)展现出强大的翻译能力,其在文档级上下文下进行APE的有效性仍不明确。本文在简单的一次性文档提示设置下,系统比较了闭源与开源LLMs的APE表现、上下文行为、鲁棒性与效率。结果表明,闭源LLMs即使仅用单次提示,也能达到接近人类水平的APE质量,且无论是否提供文档上下文均如此。这些模型对数据投毒攻击具有更高鲁棒性,但同时也暴露出无法有效利用文档级上下文进行上下文相关错误修正的局限。此外,标准自动评估指标未能可靠反映这些定性改进,凸显人工评估的必要性。尽管性能优异,闭源模型的高成本与延迟开销使其难以用于实际APE部署。总体而言,研究揭示了基于LLM的文档感知式APE的潜力与当前局限,并指向更高效长上下文建模方法的迫切需求。

原文摘要 · Abstract (English)

Automatic post-editing (APE) aims to refine machine translations by correcting residual errors. Although recent large language models (LLMs) demonstrate strong translation capabilities, their effectiveness for APE--especially under document-level context--remains insufficiently understood. We present a systematic comparison of proprietary and open-weight LLMs under a naive document-level prompting setup, analyzing APE quality, contextual behavior, robustness, and efficiency. Our results show that proprietary LLMs achieve near human-level APE quality even with simple one-shot prompting, regardless of whether document context is provided. While these models exhibit higher robustness to data poisoning attacks than open-weight counterparts, this robustness also reveals a limitation: they largely fail to exploit document-level context for contextual error correction. Furthermore, standard automatic metrics do not reliably reflect these qualitative improvements, highlighting the continued necessity of human evaluation. Despite their strong performance, the substantial cost and latency overheads of proprietary LLMs render them impractical for real-world APE deployment. Overall, our findings elucidate both the promise and current limitations of LLM-based document-aware APE, and point toward the need for more efficient long-context modeling approaches for translation refinement.

自动后编辑大模型上下文利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。