arXiv:2505.05423cs.CLcs.AI2025-05EMNLP被引 14

用专业译者反馈提升大模型对文学翻译的评估能力

LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering

  • 基于大模型问答框架,融合专业译者意见评估文学翻译质量
  • 相比现有方法,相关性提升0.07,适切性得分超顶尖指标15分以上
  • 无需参考文本,支持开源模型,适合版权敏感场景

大型语言模型在文学领域的应用日益广泛,但现有翻译评估指标仍侧重机械准确性,过度高估机器翻译而低估专业译者成果,长期可能削弱翻译质量与文化真实性。为此,我们提出LITRANSPROQA——一种参考无关、基于大模型问答的文学翻译评估框架,融入专业译者与研究者的见解,关注修辞手法、文化理解与作者风格等核心要素。实验表明,尽管文学微调的XCOMET-XL仅带来小幅提升,但LITRANSPROQA显著优于当前主流指标,在相关性上最高提升0.07,适切性评分超越最佳基线超过15分;引入译者意见作为权重进一步优化性能。其表现接近训练有素的语言学学生评估者,但仍逊于经验丰富的专业译者。该框架可应用于LLaMA3.3-70b、Qwen2.5-32b等开源模型,具备本地化处理能力,适用于受版权或伦理限制的文学翻译评估场景。

原文摘要 · Abstract (English)

The impact of Large Language Models (LLMs) has extended into literary domains. However, existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression and tend to overrate machine translation as being superior to human translation from experienced professionals. In the long run, this bias could result in an irreversible decline in translation quality and cultural authenticity. In response to the urgent need for a specialized literary evaluation metric, we introduce LITRANSPROQA, a novel, reference-free, LLM-based question-answering framework designed for literary translation evaluation. LITRANSPROQA integrates humans in the loop to incorporate insights from professional literary translators and researchers, focusing on critical elements in literary quality assessment such as literary devices, cultural understanding, and authorial voice. Our extensive evaluation shows that while literary-finetuned XCOMET-XL yields marginal gains, LITRANSPROQA substantially outperforms current metrics, achieving up to 0.07 gain in correlation and surpassing the best state-of-the-art metrics by over 15 points in adequacy assessments. Incorporating professional translator insights as weights further improves performance, highlighting the value of translator inputs. Notably, LITRANSPROQA reaches an adequacy performance comparable to trained linguistic student evaluators, though it still falls behind experienced professional translators. LITRANSPROQA shows broad applicability to open-source models like LLaMA3.3-70b and Qwen2.5-32b, indicating its potential as an accessible and training-free tool for evaluating literary translations that require local processing due to copyright or ethical considerations.

文学翻译评估指标大模型专业反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。