构建首个全球舞蹈文化理解评测基准,评估语言模型对传统舞艺的文化认知能力。
NRITYAM: Language Models Meet Art and Heritage of Dance

- 联合本土艺术家与母语者共建多语言舞蹈问答数据集
- 覆盖12种语言、9260个问题,涵盖全球12个舞蹈传统
- 适合作为文化感知类AI系统评测工具,尤其适合跨文化研究者
语言模型已成为现代工作流的关键工具,但其全球有效性依赖于对本地社会文化语境的深刻理解。为填补这一空白,我们提出NRITYAM——一个全面评估语言模型在世界舞蹈传统背景下文化理解能力的基准。该基准包含9,260个精心策划的多语言问答对,覆盖12种语言,是目前最庞大的专注于舞蹈文化知识评估的数据集。数据集由本土舞蹈艺术家和母语者从零开始协作构建,确保问题具有地域文化相关性并经严格验证。我们评估了包括大语言模型、小语言模型、多模态大模型及小多模态模型在内的多种模型。作为多语言、多文化的评测基准,NRITYAM为评估AI系统对传统表演艺术的理解与推理能力树立了新标准。详细数据样本可访问 https://github.com/niladrighosh03/NRITYAM。
原文摘要 · Abstract (English)
Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understanding of local socio-cultural contexts. To address this gap, we present NRITYAM, a comprehensive benchmark for evaluating the cultural comprehension capabilities of language models in the context of global dance traditions. NRITYAM comprises 9,260 carefully curated question-answer pairs spanning 12 languages, making it the largest dataset dedicated to evaluating cultural knowledge in dance. The dataset has been developed from the ground up through close collaboration with native dance artists and native speakers of the languages, who authored and validated culturally relevant questions specific to their regions. We evaluate a broad set of models, including large language models, small language models, multimodal large language models, and small multimodal language models. As a multilingual and multicultural benchmark, NRITYAM sets a new standard for evaluating the ability of AI systems to understand and reason about traditional performing arts. Detailed dataset samples are available at~\url{https://github.com/niladrighosh03/NRITYAM}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。