评测大模型对波斯谚语的文化理解能力,发现跨文化翻译性能不足。
MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
- 构建波斯谚语多维度评测集,测试上下文与跨文化理解。
- 模型识别波斯谚语准确率超90%,但翻译成英语谚语仅达79%。
- 揭示大模型在文化知识与类比推理上的短板,适合低资源语言研究者。
近年来,多语言大语言模型已成为日常生活中不可或缺的一部分,使其掌握对话语言规则以实现有效沟通至关重要。尽管已有研究评估了大模型在高资源语言中对隐喻语言的理解能力,但在低资源语言中的表现仍缺乏探索。本文提出MasalBench,一个全面的基准测试,用于评估大模型对波斯谚语的上下文与跨文化理解能力,这些谚语是该低资源语言对话中的关键组成部分。我们在八个最先进的大模型上评估了MasalBench,发现它们在识别上下文中的波斯谚语方面表现良好,准确率均高于0.90。然而,当任务转换为识别对应的英文谚语时,性能显著下降,最佳模型准确率为0.79。研究结果突显了当前大模型在文化知识和类比推理方面的局限性,并为其他低资源语言的跨文化理解评估提供了框架。MasalBench可在https://github.com/kalhorghazal/MasalBench获取。
原文摘要 · Abstract (English)
In recent years, multilingual Large Language Models (LLMs) have become an inseparable part of daily life, making it crucial for them to master the rules of conversational language in order to communicate effectively with users. While previous work has evaluated LLMs' understanding of figurative language in high-resource languages, their performance in low-resource languages remains underexplored. In this paper, we introduce MasalBench, a comprehensive benchmark for assessing LLMs' contextual and cross-cultural understanding of Persian proverbs, which are a key component of conversation in this low-resource language. We evaluate eight state-of-the-art LLMs on MasalBench and find that they perform well in identifying Persian proverbs in context, achieving accuracies above 0.90. However, their performance drops considerably when tasked with identifying equivalent English proverbs, with the best model achieving 0.79 accuracy. Our findings highlight the limitations of current LLMs in cultural knowledge and analogical reasoning, and they provide a framework for assessing cross-cultural understanding in other low-resource languages. MasalBench is available at https://github.com/kalhorghazal/MasalBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。