arXiv:2605.02601cs.CL2026-05ACL被引 1

评测大模型在30+语言文化下的常识理解能力,聚焦低资源语言。

SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

论文配图:SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
图 1 · 摘自论文原文
  • 基于扩展的BLEnD基准,覆盖30多个语言-文化对。
  • 仅允许评估,禁止用数据训练或微调模型。
  • 吸引62支队伍参赛,揭示低资源语言中的模型偏差问题。

我们提出了一个评估大语言模型和NLP系统在多语言、多文化环境下适应能力的共享任务。任务数据基于我们手工构建的BLEnD基准(Myung et al. 2024)扩展而来,涵盖超过30个语言-文化组合,主要代表全球多个大陆上的低资源语言。任务严格限定为评估用途,参与者不得使用数据进行训练、微调、少样本学习或任何其他形式的模型修改。任务包含两个赛道:(a) 短答案问答(SAQ)和 (b) 多选题(MCQ)。参与者需预测标签,并可自由使用任意NLP系统及建模策略,但必须确保基准仅用于评估。本任务吸引了140余名注册参与者,最终收到62支团队的提交结果,以及19篇系统描述论文。我们报告了实验结果,并分析了表现最佳系统与最常用方法。此外,还讨论了评估中的共性挑战、模型错位现象,以及在低资源语言和代表性不足文化背景下模型行为的方法论思考。

原文摘要 · Abstract (English)

We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al. 2024), covering more than 30 language-culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification. Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers. We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures.

多语言低资源常识理解评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。