构建首个覆盖全印度文化的语言模型评测基准
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
- 设计21,853个跨28省的文化问答对,覆盖16类文化维度
- 主流大模型在区域文化题上准确率普遍偏低,本地化能力弱
- 适合研究本土化AI、文化认知或南亚语言模型的学者使用
语言模型虽广泛应用于现代工作流,但其全球有效性依赖对本地社会文化背景的理解。为此,我们提出SANSKRITI,一个评估语言模型对印度文化理解力的综合性基准。该数据集包含21,853个精心策划的问答对,覆盖28个州和8个联邦属地,是目前最大的印度文化知识测评数据集。涵盖十六项核心文化属性:仪式与典礼、历史、旅游、饮食、舞蹈与音乐、服饰、语言、艺术、节日、宗教、医学、交通、体育、夜生活及人物,全面呈现印度文化多样性。我们在领先的大型语言模型(LLMs)、印地语语言模型(ILMs)和小型语言模型(SLMs)上评估该基准,发现模型在处理文化细微差别时存在显著差异,尤其在区域特定问题上表现不佳。SANSKRITI通过提供广泛、丰富且多元的数据,为评估和提升语言模型的文化理解能力设立了新标准。
原文摘要 · Abstract (English)
Language Models (LMs) are indispensable tools shaping modern workflows, but their global effectiveness depends on understanding local socio-cultural contexts. To address this, we introduce SANSKRITI, a benchmark designed to evaluate language models' comprehension of India's rich cultural diversity. Comprising 21,853 meticulously curated question-answer pairs spanning 28 states and 8 union territories, SANSKRITI is the largest dataset for testing Indian cultural knowledge. It covers sixteen key attributes of Indian culture: rituals and ceremonies, history, tourism, cuisine, dance and music, costume, language, art, festivals, religion, medicine, transport, sports, nightlife, and personalities, providing a comprehensive representation of India's cultural tapestry. We evaluate SANSKRITI on leading Large Language Models (LLMs), Indic Language Models (ILMs), and Small Language Models (SLMs), revealing significant disparities in their ability to handle culturally nuanced queries, with many models struggling in region-specific contexts. By offering an extensive, culturally rich, and diverse dataset, SANSKRITI sets a new standard for assessing and improving the cultural understanding of LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。