评测大模型在中英诗歌翻译中的意图保留能力,发现其存在文化失真问题。
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
- 构建包含科技、历史与文学的多类型语料库,用回译+弗里德曼检验评估性能
- 大模型在科学文本中表现尚可,但在文学翻译中难以保留文化意境
- 揭示'诗意意图悖论'现象,适合关注AI翻译文化保真的研究者阅读
大语言模型(LLMs)的快速发展重塑了机器翻译格局,但在保留中文翻译中的诗意意图、文化传承及专业术语方面仍面临挑战。本研究构建了一个涵盖中文科技术语、历史翻译悖论和文学隐喻的多样化语料库。采用基于回译与弗里德曼检验的评估系统(BT-Fried),对六种主流大模型(如GPT-4.5、DeepSeek V3)和三种传统翻译工具进行评估,指标包括BLEU、CHRF、TER及语义相似度。关键发现:(1) 科学摘要常因回译获益,而传统工具在语言差异显著的文本中优于大模型;(2) 大模型在文化与文学内容保留上表现不佳,体现‘诗意意图悖论’;(3) 部分模型出现‘逐字回译’现象,反映其潜在的记忆行为;(4) 提出一种使用结巴分词与n-gram加权的新型BLEU变体。研究为中文自然语言处理性能提供了实证评估,并深化了对AI翻译中文化忠实性的理解。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has reshaped the landscape of machine translation, yet challenges persist in preserving poetic intent, cultural heritage, and handling specialized terminology in Chinese-English translation. This study constructs a diverse corpus encompassing Chinese scientific terminology, historical translation paradoxes, and literary metaphors. Utilizing a back-translation and Friedman test-based evaluation system (BT-Fried), we evaluate BLEU, CHRF, TER, and semantic similarity metrics across six major LLMs (e.g., GPT-4.5, DeepSeek V3) and three traditional translation tools. Key findings include: (1) Scientific abstracts often benefit from back-translation, while traditional tools outperform LLMs in linguistically distinct texts; (2) LLMs struggle with cultural and literary retention, exemplifying the "paradox of poetic intent"; (3) Some models exhibit "verbatim back-translation", reflecting emergent memory behavior; (4) A novel BLEU variant using Jieba segmentation and n-gram weighting is proposed. The study contributes to the empirical evaluation of Chinese NLP performance and advances understanding of cultural fidelity in AI-mediated translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。