对比人类与机器翻译的直译程度,发现大模型能自我优化为更自然表达。
Testing the Deliteralization Hypothesis in Human and Machine Translation

- 用六种启发式方法构建量化直译度指标,评估不同系统输出
- 大模型在自迭代修订中直译度持续下降,首次验证该现象天然存在于生成过程
- 作为校对者时,大模型偏好修改口语化表达而非直译内容,与人类相反
近期从专用神经机器翻译系统转向通用大语言模型(LLM)重塑了机器翻译范式,已有研究指出大模型生成结果更流畅、更少直译。本文通过WMT24++数据集,比较54种语言对下两种NMT系统和六种大模型在直接翻译、迭代自修订及人工初稿后编辑任务中的直译程度。使用基于六种启发式的有效合成直译指数进行测量。结果表明:(i) 人类译文仍显著低于所有测试的机器翻译系统,尽管近年大模型缩小了差距;(ii) 当被要求迭代修订自身输出时,大模型表现出单调降直译趋势,首次证明该现象可原生存在于大模型生成中;(iii) 作为后编辑器时,大模型容忍直译初稿并聚焦于改写符合人类语感的表达,与人类后编辑者的修正逻辑相反。
原文摘要 · Abstract (English)
The recent shift from dedicated NMT systems to general-purpose LLMs has reshaped machine translation, with LLMs reported to produce more fluent, less literal output than their predecessors. We test whether this shift extends to the deliteralization hypothesis, the long-standing claim from translation studies that translations become progressively less literal as they are drafted and revised. Using the WMT24++ dataset, we compare the literality of human translations and post-editions to that of two NMT systems and six LLMs across 54 language pairs and three tasks: direct translation, iterative self-revision, and post-editing of human drafts. Literality is measured via a validated Synthetic Literality Index built from six heuristics. We find that (i) human translations remain significantly less literal than those of all tested MT systems, though recent LLMs narrow the gap; (ii) when prompted to iteratively revise their own output, LLMs deliteralize monotonically, providing the first evidence that the hypothesis applies natively to LLM generation; and (iii) as post-editors, LLMs invert the revision triggers of human post-editors, tolerating literal drafts and targeting idiomatic human formulations for revision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。