系统研究大模型翻译优化,发现分段精炼效果最佳。
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation

- 文档级翻译后接段落级精炼,提升最稳定
- 精炼主要改善流畅度、风格和术语,对准确率提升有限
- 通用提示比针对性修复更有效,适合多数场景
迭代自精炼是一种简单的推理阶段策略:大模型在多次推理中自我修正翻译结果。然而,文档级精炼仍缺乏深入理解:1)哪种流程最优,2)哪些质量维度提升,3)精炼器行为如何。本文针对文档级文学翻译进行系统研究,涵盖九种大模型与七种语言对。在九种精炼粒度组合与五种精炼策略下,发现稳健方案:先进行文档级翻译,再进行段落级精炼,能带来强且稳定的提升。相反,文档级精炼常修改较少,收益更小且不可靠。除粒度外,简单通用精炼提示始终优于错误特定提示和评估-修复方案。大规模人工评估显示,精炼增益主要来自流畅性、风格和术语,准确性提升有限且不一致。模型强度实验表明,精炼会将输出推向精炼器的分布,而非精准修复错误。这些发现揭示了当前精炼方法的机制与局限。
原文摘要 · Abstract (English)
Iterative self-refinement is a simple inference-time strategy for machine translation: an LLM revises its own translation over multiple inference-time passes. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering nine LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, we find a robust recipe: document-level MT followed by segment-level refinement yields strong and stable improvements. In contrast, document-level refinement often makes fewer edits and leads to smaller or less reliable gains. Beyond granularity, A simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. Our large-scale human evaluation shows that refinement gains come primarily from fluency, style, and terminology, with limited and less consistent improvements in adequacy. Experiments varying model strength reveal refinement projects outputs toward the refiner's distribution rather than performing targeted error repair. These findings clarify the mechanisms and limitations of current refinement approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。