对比不同提示方法与模型组合,提升波斯语与英语习语翻译质量。
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
- 构建双语习语平行数据集,覆盖2200个波斯语习语及用例。
- Claude-3.5-Sonnet在双向翻译中表现最佳,准确率显著领先。
- 结合弱模型与Google Translate可提升波斯语翻译效果。
大型语言模型(LLMs)在翻译隐喻语言方面优于神经机器翻译(NMT)系统。然而,不同提示方法及LLM-NMT组合对习语翻译的影响尚未充分研究。本文构建了两个平行语料库,用于波斯语↔英语习语翻译,其中波斯语习语源自包含2200个习语及其释义的PersianIdioms资源,700个含使用示例。基于此,我们评估多种开源与闭源LLM、NMT模型及其组合,通过习语翻译准确率与流畅性评估性能。自动评价方法如LLM-as-a-judge、BLEU和BERTScore能有效比较模型表现。实验表明,Claude-3.5-Sonnet在双向翻译中均表现优异;英语→波斯语翻译中,弱LLM与Google Translate结合效果更佳;波斯语→英语翻译中,简单模型适用单提示,高级模型则需复杂提示。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown superior capabilities in translating figurative language compared to neural machine translation (NMT) systems. However, the impact of different prompting methods and LLM-NMT combinations on idiom translation has yet to be thoroughly investigated. This paper introduces two parallel datasets of sentences containing idiomatic expressions for Persian$\rightarrow$English and English$\rightarrow$Persian translations, with Persian idioms sampled from our PersianIdioms resource, a collection of 2,200 idioms and their meanings, with 700 including usage examples. Using these datasets, we evaluate various open- and closed-source LLMs, NMT models, and their combinations. Translation quality is assessed through idiom translation accuracy and fluency. We also find that automatic evaluation methods like LLM-as-a-judge, BLEU, and BERTScore are effective for comparing different aspects of model performance. Our experiments reveal that Claude-3.5-Sonnet delivers outstanding results in both translation directions. For English$\rightarrow$Persian, combining weaker LLMs with Google Translate improves results, while Persian$\rightarrow$English translations benefit from single prompts for simpler models and complex prompts for advanced ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。