arXiv:2508.10421cs.CL2025-08中稿 · COLM被引 2

评测大模型中文成语翻译,发现主流系统错误率超两成

Evaluating LLMs on Chinese Idiom Translation

  • 构建成语翻译错误分类框架,系统分析九个模型表现
  • GPT-4最佳仍错28%,多数翻译出现字面误译或缺失
  • 现有评估指标相关性不足0.48,提出新检测模型提升准确率

成语因其隐喻意义常与字面意思不同,在中文中尤为常见,多含历史典故且有固定结构。尽管大语言模型在机器翻译上取得进展,但中文成语翻译仍缺乏系统研究。本文提出IdiomEval框架,包含全面的错误分类体系,并人工标注了来自九个现代系统的900组翻译对,涵盖网页、新闻、维基百科和社交媒体四类领域。结果显示,这些系统在成语翻译中普遍表现不佳,常出现错误、字面化、部分或遗漏翻译。表现最好的GPT-4在28%的情况下出错。此外,现有评估指标与人工评分的皮尔逊相关性低于0.48,效果有限。为此,我们开发了改进模型,实现成语翻译错误检测F₁得分为0.68。

原文摘要 · Abstract (English)

Idioms, whose figurative meanings usually differ from their literal interpretations, are common in everyday language, especially in Chinese, where they often contain historical references and follow specific structural patterns. Despite recent progress in machine translation with large language models, little is known about Chinese idiom translation. In this work, we introduce IdiomEval, a framework with a comprehensive error taxonomy for Chinese idiom translation. We annotate 900 translation pairs from nine modern systems, including GPT-4o and Google Translate, across four domains: web, news, Wikipedia, and social media. We find these systems fail at idiom translation, producing incorrect, literal, partial, or even missing translations. The best-performing system, GPT-4, makes errors in 28% of cases. We also find that existing evaluation metrics measure idiom quality poorly with Pearson correlation below 0.48 with human ratings. We thus develop improved models that achieve F$_1$ scores of 0.68 for detecting idiom translation errors.

成语翻译大模型评测语言理解机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。