评测大模型对孟加拉文化知识的掌握,发现加上下文后表现大幅提升。
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
- 构建孟加拉文化知识数据集,涵盖民俗、烹饪和方言
- 无上下文时模型文化知识准确率低,加上下文后显著提升
- 适合研究低资源文化认知与多语言模型改进的学者
近年来自然语言处理研究展示了大语言模型在多种任务上的强大能力。尽管已有多种多语言基准推动了模型的文化评估,但在低资源文化细节捕捉方面仍存在明显不足。本文通过构建涵盖民俗传统、烹饪艺术和地方方言的孟加拉语文化知识(BLanCK)数据集,填补这一空白。对多个多语言模型的评估显示,这些模型在非文化类任务中表现良好,但在文化知识任务上显著落后;而提供上下文信息后,所有模型性能均大幅提升,凸显上下文感知架构与文化定制训练数据的重要性。
原文摘要 · Abstract (English)
Recent progress in NLP research has demonstrated remarkable capabilities of large language models (LLMs) across a wide range of tasks. While recent multilingual benchmarks have advanced cultural evaluation for LLMs, critical gaps remain in capturing the nuances of low-resource cultures. Our work addresses these limitations through a Bengali Language Cultural Knowledge (BLanCK) dataset including folk traditions, culinary arts, and regional dialects. Our investigation of several multilingual language models shows that while these models perform well in non-cultural categories, they struggle significantly with cultural knowledge and performance improves substantially across all models when context is provided, emphasizing context-aware architectures and culturally curated training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。