构建马来文化评测基准,检验大模型在低资源语言下的文化理解能力
MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints
- 设计开放式多选题格式,减少猜测偏差,提升评估公平性
- 跨6大文化领域测试,发现主流大模型在马来语文化理解上表现参差
- 适合关注AI文化偏见、多语言公平性的研究者与开发者
大型语言模型因训练数据主要来自英语和中文等高资源语言,常表现出文化偏见,难以准确呈现和评估多元文化背景,尤其在低资源语言环境下。为此,我们提出MyCulture,一个针对马来西亚文化的综合性评测基准,涵盖艺术、服饰、习俗、娱乐、食物和宗教六大支柱,全部以马来语呈现。不同于传统评测,MyCulture采用新颖的开放式多选题形式,无预设选项,有效减少猜测行为并缓解格式偏差。我们从理论上论证了该结构在提升评估公平性与区分度方面的有效性。通过对比结构化与自由输出的表现,分析模型在不同语言提示下的表现差异,我们评估了多个区域及国际主流大模型,结果揭示其在文化理解方面存在显著差距,凸显出在大模型研发与评估中建立文化扎根、语言包容性评测体系的紧迫性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cultural contexts, particularly in low-resource language settings. To address this, we introduce MyCulture, a benchmark designed to comprehensively evaluate LLMs on Malaysian culture across six pillars: arts, attire, customs, entertainment, food, and religion presented in Bahasa Melayu. Unlike conventional benchmarks, MyCulture employs a novel open-ended multiple-choice question format without predefined options, thereby reducing guessing and mitigating format bias. We provide a theoretical justification for the effectiveness of this open-ended structure in improving both fairness and discriminative power. Furthermore, we analyze structural bias by comparing model performance on structured versus free-form outputs, and assess language bias through multilingual prompt variations. Our evaluation across a range of regional and international LLMs reveals significant disparities in cultural comprehension, highlighting the urgent need for culturally grounded and linguistically inclusive benchmarks in the development and assessment of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。