arXiv:2604.15165cs.CL2026-04中稿 · _papers

大模型翻译时会过度生成,需区分其是胡编乱造还是合理解释。

Fabricator or dynamic translator?

论文配图:Fabricator or dynamic translator?
图 1 · 摘自论文原文
  • 通过分析生成内容判断是否为过度补充,区分胡说与恰当说明。
  • 在真实场景中发现模型会自动生成解释或虚构细节,影响翻译可信度。
  • 适合关注AI翻译质量与可信度的研究者和产品经理。

大语言模型在机器翻译中表现出色,但因其生成特性可能产生过度生成现象。这些过度生成不同于神经机器翻译中的神经胡言,表现为模型自我解释、冒险虚构或恰当地提供说明,使其行为类似人类翻译者,提升目标受众的理解。检测并确定过度生成的具体性质是一项挑战。本文详述了在商业环境中探索的不同策略及其结果。

原文摘要 · Abstract (English)

LLMs are proving to be adept at machine translation although due to their generative nature they may at times overgenerate in various ways. These overgenerations are different from the neurobabble seen in NMT and range from LLM self-explanations, to risky confabulations, to appropriate explanations, where the LLM is able to act as a human translator would, enabling greater comprehension for the target audience. Detecting and determining the exact nature of the overgenerations is a challenging task. We detail different strategies we have explored for our work in a commercial setting, and present our results.

大模型机器翻译过度生成可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。