用自动化框架评估大模型中英翻译,发现文学类文本仍存挑战
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
- 构建基于语义情感分析的自动化评估框架
- 大模型在新闻翻译中表现优异,文学翻译差异明显
- 深度求索在文化细节和语法还原上更优,但古典表达仍难处理
尽管大型语言模型(LLMs)在机器翻译中表现卓越,但对其翻译质量的系统性评估仍有限。核心挑战在于缺乏自动化评估框架,人工评测耗时且难以覆盖快速演进的模型与多样文本。本文提出一种基于语义与情感分析的自动化机器学习框架,评估谷歌翻译及GPT-4、GPT-4o、DeepSeek等模型在中英翻译中的表现。对比了涵盖现代与古典文学、新闻报道等高影响力中文文本的原文与译文。采用新型相似度指标衡量翻译质量,并由专家译者进行人工验证。结果表明:大模型在新闻翻译中表现良好,但在文学文本上表现分化;尽管GPT-4o与DeepSeek在复杂语境下语义保留更佳,但DeepSeek在文化细微差别与语法呈现方面更具优势。然而,文化细节、古典引用与修辞表达的精准传递仍是所有模型面临的开放问题。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) have exceptional performance in machine translation, only a limited systematic assessment of translation quality has been done. The challenge lies in automated frameworks, as human-expert-based evaluations can be time-consuming, given the fast-evolving LLMs and the need for a diverse set of texts to ensure fair assessments of translation quality. In this paper, we utilise an automated machine learning framework featuring semantic and sentiment analysis to assess Mandarin Chinese to English translation using Google Translate and LLMs, including GPT-4, GPT-4o, and DeepSeek. We compare original and translated texts in various classes of high-profile Chinese texts, which include novel texts that span modern and classical literature, as well as news articles. As the main evaluation measures, we utilise novel similarity metrics to compare the quality of translations produced by LLMs and further evaluate them by an expert human translator. Our results indicate that the LLMs perform well in news media translation, but show divergence in their performance when applied to literary texts. Although GPT-4o and DeepSeek demonstrated better semantic conservation in complex situations, DeepSeek demonstrated better performance in preserving cultural subtleties and grammatical rendering. Nevertheless, the subtle challenges in translation remain: maintaining cultural details, classical references and figurative expressions remain an open problem for all the models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。