arXiv:2412.09993cs.CL2024-12

对比不同提示方法与模型组合,提升波斯语与英语习语翻译质量。

Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation

  • 构建双语习语平行数据集,覆盖2200个波斯语习语及用例。
  • Claude-3.5-Sonnet在双向翻译中表现最佳,准确率显著领先。
  • 结合弱模型与Google Translate可提升波斯语翻译效果。

大型语言模型(LLMs)在翻译隐喻语言方面优于神经机器翻译(NMT)系统。然而,不同提示方法及LLM-NMT组合对习语翻译的影响尚未充分研究。本文构建了两个平行语料库,用于波斯语↔英语习语翻译,其中波斯语习语源自包含2200个习语及其释义的PersianIdioms资源,700个含使用示例。基于此,我们评估多种开源与闭源LLM、NMT模型及其组合,通过习语翻译准确率与流畅性评估性能。自动评价方法如LLM-as-a-judge、BLEU和BERTScore能有效比较模型表现。实验表明,Claude-3.5-Sonnet在双向翻译中均表现优异;英语→波斯语翻译中,弱LLM与Google Translate结合效果更佳;波斯语→英语翻译中,简单模型适用单提示,高级模型则需复杂提示。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom translation has yet to be thoroughly investigated. This paper introduces two parallel datasets of sentences containing idiomatic expressions for Persian$\rightarrow$English and English$\rightarrow$Persian translations, with Persian idioms sampled from our PersianIdioms resource, a collection of 2,200 idioms and their meanings, with 700 including usage examples. Using these datasets, we evaluate various open- and closed-source LLMs, NMT models, and their combinations. Translation quality is assessed through idiom translation accuracy and fluency. We also find that automatic evaluation methods like LLM-as-a-judge, BLEU, and BERTScore are effective for comparing different aspects of model performance. Our experiments reveal that Claude-3.5-Sonnet delivers outstanding results in both translation directions. For English$\rightarrow$Persian, combining weaker LLMs with Google Translate improves results, while Persian$\rightarrow$English translations benefit from single prompts for simpler models and complex prompts for advanced ones.

习语翻译多语言LLM波斯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。