首个聚焦孟加拉本土文化与语言的评测基准,测试大模型在孟加拉语理解中的表现。
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
- 构建2366道多选题,覆盖孟加拉历史、文化与语言23个类别。
- 主流大模型在孟加拉语音学上表现不佳,整体水平低于英语等主流语言。
- 适合研究低资源语言、文化敏感性模型或孟加拉语AI应用的开发者使用。
本文提出BLUCK,一个用于评估大语言模型在孟加拉语语言理解与文化知识方面表现的新数据集。该数据集包含2366道精心筛选的多项选择题(MCQs),源自多个大学及求职考试的合集,涵盖23个类别,涉及孟加拉国文化、历史及孟加拉语语言学内容。我们使用6个专有和3个开源大模型(包括GPT-4o、Claude-3.5-Sonnet、Gemini-1.5-Pro、Llama-3.3-70B-Instruct和DeepSeekV3)对BLUCK进行了基准测试。结果显示,尽管模型整体表现尚可,但在孟加拉语语音学方面仍存在明显短板。当前大模型在孟加拉语文化与语言理解方面的性能尚未达到英语等主流语言水平,但表明孟加拉语处于中等资源语言地位。尤为重要的是,BLUCK是首个以原生孟加拉文化、历史与语言学为核心的多选题评测基准。
原文摘要 · Abstract (English)
In this work, we introduce BLUCK, a new dataset designed to measure the performance of Large Language Models (LLMs) in Bengali linguistic understanding and cultural knowledge. Our dataset comprises 2366 multiple-choice questions (MCQs) carefully curated from compiled collections of several college and job level examinations and spans 23 categories covering knowledge on Bangladesh's culture and history and Bengali linguistics. We benchmarked BLUCK using 6 proprietary and 3 open-source LLMs - including GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, Llama-3.3-70B-Instruct, and DeepSeekV3. Our results show that while these models perform reasonably well overall, they, however, struggles in some areas of Bengali phonetics. Although current LLMs' performance on Bengali cultural and linguistic contexts is still not comparable to that of mainstream languages like English, our results indicate Bengali's status as a mid-resource language. Importantly, BLUCK is also the first MCQ-based evaluation benchmark that is centered around native Bengali culture, history, and linguistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。