arXiv:2510.01247cs.CLcs.AI2025-10EMNLP被引 2

构建跨文化体育知识测试集,评估大模型对全球传统运动的理解能力

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

  • 设计多模态跨文化体育问答数据集,覆盖60国4类文化
  • 3.3万道题目涵盖历史、规则与场景三类,支持零样本/少样本推理
  • 首次系统评估大模型对非主流体育的认知,适合研究AI文化理解的学者

语言模型主要在流行体育上进行评估,常忽略地区性和原住民体育传统。为填补这一空白,我们提出 extbf{ extit{CultSportQA}},一个用于评估语言模型对60个国家、6大洲传统体育理解能力的基准测试。该数据集包含3.3万道多模态(文本与图像)选择题,按历史、规则、场景三类划分。我们采用零样本、少样本及思维链(CoT)提示法,在多种大语言模型(LLMs)、小语言模型(SLMs)和多模态大语言模型(MLMs)上进行评估。 extbf{ extit{CultSportQA}}为评估AI对传统体育的理解与推理能力提供了全面的多语言、多文化标准。

原文摘要 · Abstract (English)

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce \textbf{\textit{CultSportQA}}, a benchmark designed to assess LMs' understanding of traditional sports across 60 countries and 6 continents, encompassing four distinct cultural categories. The dataset features 33,000 multiple-choice questions (MCQs) across text and image modalities, each of which is categorized into three key types: history-based, rule-based, and scenario-based. To evaluate model performance, we employ zero-shot, few-shot, and chain-of-thought (CoT) prompting across a diverse set of Large Language Models (LLMs), Small Language Models (SLMs), and Multimodal Large Language Models (MLMs). By providing a comprehensive multilingual and multicultural sports benchmark, \textbf{\textit{CultSportQA}} establishes a new standard for assessing AI's ability to understand and reason about traditional sports.

跨文化理解体育认知多模态评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。