arXiv:2604.18169cs.CLcs.AI2026-04ACL被引 1

评测大模型在文学翻译中理解与创造的协同能力

Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation

论文配图:Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation
图 1 · 摘自论文原文
  • 设计配对任务框架,同步评估原文理解与翻译创造力
  • 仅1个模型(Mistral-Large)接近人类创造力水平,多数模型创造力得分低于0.1
  • 中文-英文翻译差距显著,创意提示效果有限

大语言模型在文学翻译等创造性任务中的应用日益广泛,但译文创造力仍缺乏系统评估,且原文理解常被孤立研究。本文提出一种配对任务框架,基于11本书籍的文学节选,任务一评估源文理解,任务二通过创意潜力单元(UCPs,如隐喻、双关语)衡量翻译创造力。采用结合专家标注与基于UCP的自动评分的可扩展评估方案,对23个模型和4种创意导向提示进行基准测试。结果表明,强理解力并不等同于高水平创造力:模型多产出直译或语境不符的译文,尤其在英语-中文这对语言间差距更大。创意提示仅带来小幅提升,仅一个模型(Mistral-Large)接近人类水平(0.167 vs. 0.246)。所有模型-提示组合中,仅三个创造力得分超过0.1,其余均接近零。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for creative tasks such as literary translation. Yet translational creativity remains underexplored and is rarely evaluated at scale, while source-text comprehension is typically studied in isolation, despite the fact that, in professional translation, comprehension and creativity are tightly intertwined. We address these gaps with a paired-task framework applied to literary excerpts from 11 books. Task 1 assesses source-text comprehension, and Task 2 evaluates translational creativity through Units of Creative Potential (UCPs), such as metaphors and wordplay. Using a scalable evaluation setup that combines expert human annotations with UCP-based automatic scoring, we benchmark 23 models and four creativity-oriented prompts. Our findings show that strong comprehension does not translate into human-level creativity: models often produce literal or contextually inappropriate renderings, with particularly large gaps for the more distant English-Chinese language pair. Creativity-oriented prompts yield only modest gains, and only one model, Mistral-Large, comes close to human-level creativity (0.167 vs. 0.246). Across all model-prompt combinations, only three exceed a creativity score of 0.1, while the rest remain at or near zero.

文学翻译创造力评测大模型评估跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。