arXiv:2410.16461cs.CL2024-10被引 15

对比大模型在多语言隐喻和成语理解上的表现差异。

Comparative Study of Multilingual Idioms and Similes in Large Language Models

  • 用链式思考、少样本等提示策略评估多语言图文理解能力
  • 开放模型在低资源语言的比喻理解上表现较差
  • 成语理解已接近饱和,需更难评测数据

本研究填补了现有文献中关于大语言模型在多语言环境下对不同修辞类型(比喻与成语)理解能力比较的空白。通过两个多语言数据集对模型进行评估,并引入链式思考、少样本及英文翻译提示等策略。我们还将数据集扩展至波斯语,构建了两个新评估集。综合评估涵盖闭源模型(GPT-3.5、GPT-4o mini、Gemini 1.5)与开源模型(Llama 3.1、Qwen2),揭示了不同语言、修辞类型与模型间的显著性能差异。结果显示,提示工程总体有效,但效果因修辞类型、语言和模型而异;开源模型在低资源语言的比喻理解上尤其薄弱。此外,多数语言的成语理解已趋于饱和,亟需更具挑战性的评测标准。

原文摘要 · Abstract (English)

This study addresses the gap in the literature concerning the comparative performance of LLMs in interpreting different types of figurative language across multiple languages. By evaluating LLMs using two multilingual datasets on simile and idiom interpretation, we explore the effectiveness of various prompt engineering strategies, including chain-of-thought, few-shot, and English translation prompts. We extend the language of these datasets to Persian as well by building two new evaluation sets. Our comprehensive assessment involves both closed-source (GPT-3.5, GPT-4o mini, Gemini 1.5), and open-source models (Llama 3.1, Qwen2), highlighting significant differences in performance across languages and figurative types. Our findings reveal that while prompt engineering methods are generally effective, their success varies by figurative type, language, and model. We also observe that open-source models struggle particularly with low-resource languages in similes. Additionally, idiom interpretation is nearing saturation for many languages, necessitating more challenging evaluations.

多语言修辞理解大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。