自动评估工具常误判文学翻译中的创意表达,偏向机器译文。
Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

- 用专业译者标注的多语言文学译文数据集,评估自动评分效果。
- 自动评分与专家评价相关性低,尤其对诗歌等文学体裁更差。
- 大模型评分会系统性偏爱机器译文,打压文化适配的创造性改写。
本文研究了自动评估指标(AEMs)和大模型作为评判者(LLM-as-a-judge)在多种语言、文体及翻译方式下的文学翻译评价表现,旨在评估其与专业人士评价的一致性,以及是否可替代人工标注。研究构建了一个涵盖三种翻译模式(人工译、机器译、后编辑)、三种文体和三种语对的文学翻译数据集,并由资深专业译者详细标注了创造力维度。结果表明,两类自动评估方法在创造力评价上与专家评价相关性均较差,且大模型评委会系统性偏好机器译文,惩罚具有创造性和文化适配性的翻译方案。尤其在诗歌等高度文学化文体中表现更差,揭示了当前自动评估工具在文学翻译领域的根本局限,亟需开发不将非常规表达误判为错误的新评估体系。
原文摘要 · Abstract (English)
This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess how well these tools align with professionals when evaluating translation, creativity (creative shifts & errors), and see if they can substitute laborious manual annotations. A dataset of literary translations across three modalities (human translation, machine translation, and post-editing), three genres and three language pairs was created and annotated in detail for creativity by experienced professional literary translators. The results show that both AEMs and LLM-as-a-judge evaluations correlate poorly with professional evaluations on creativity, with LLM-as-a-judge showing a systematic bias in favour of machine-translated texts and penalising creative and culturally appropriate solutions. Moreover, performance is consistently worse for more literary genres such as poetry. This highlights fundamental limitations of current automatic evaluation tools for literary translation and the need to create new tools that do not frequently consider out of routine translations as errors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。