arXiv:2412.01340cs.CL2024-12被引 1

提出两阶段框架,细粒度评估英译韩文学翻译质量。

A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls

  • 分两步评估文学翻译,提升可解释性
  • 与人工判断相关性高于传统指标
  • 适合关注文化敏感性的翻译研究者

本文提出并评估了一种两阶段流程,用于从英文到韩文的文学机器翻译进行细粒度评估。结果表明,该框架提供的度量标准更具细粒度且可解释,与人工判断的相关性高于传统机器翻译评价指标。然而,在韩语敬语等维度上仍未能达到人与人之间的判断一致性。我们还发现大型语言模型倾向于偏好其他大模型生成的翻译,强调需发展更精细的评估方法,以确保文学翻译在文化和语义上的准确传达。

原文摘要 · Abstract (English)

In this work, we propose and evaluate the feasibility of a two-stage pipeline to evaluate literary machine translation, in a fine-grained manner, from English to Korean. The results show that our framework provides fine-grained, interpretable metrics suited for literary translation and obtains a higher correlation with human judgment than traditional machine translation metrics. Nonetheless, it still fails to match inter-human agreement, especially in metrics like Korean Honorifics. We also observe that LLMs tend to favor translations generated by other LLMs, and we highlight the necessity of developing more sophisticated evaluation methods to ensure accurate and culturally sensitive machine translation of literary works.

文学翻译评估框架LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。