评测大模型翻译谚语的能力,发现文化背景影响翻译效果。
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
- 构建四语对谚语独立与对话中的翻译数据集
- 大模型在相似文化语言间翻译表现更好,优于传统模型
- 现有评估指标难以可靠衡量谚语翻译质量
尽管机器翻译取得了显著进展,但针对语言中文化元素(如习语、谚语和口语表达)的翻译研究仍不充分。本文探究了先进神经机器翻译(NMT)与大语言模型(LLMs)在谚语翻译方面的能力。我们为四种语言对构建了独立谚语及对话中谚语的翻译数据集。实验表明,这些模型在具有相似文化背景的语言间可实现良好翻译,且大模型整体表现优于传统NMT模型。此外,我们发现当前自动评估指标(如BLEU、CHRF++、COMET)在评估谚语翻译质量时存在不足,凸显了开发更具文化敏感性的评估方法的必要性。
原文摘要 · Abstract (English)
Despite achieving remarkable performance, machine translation (MT) research remains underexplored in terms of translating cultural elements in languages, such as idioms, proverbs, and colloquial expressions. This paper investigates the capability of state-of-the-art neural machine translation (NMT) and large language models (LLMs) in translating proverbs, which are deeply rooted in cultural contexts. We construct a translation dataset of standalone proverbs and proverbs in conversation for four language pairs. Our experiments show that the studied models can achieve good translation between languages with similar cultural backgrounds, and LLMs generally outperform NMT models in proverb translation. Furthermore, we find that current automatic evaluation metrics such as BLEU, CHRF++ and COMET are inadequate for reliably assessing the quality of proverb translation, highlighting the need for more culturally aware evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。