构建首个覆盖45个地区的跨文化知识测试集,挑战大模型真实文化理解能力。
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- 通过人机协作方式生成并验证1696道文化题,覆盖全球45个地区。
- 顶尖模型在难版测试中准确率仅28.7%~61.5%,远低于人类92.4%。
- 发现模型易误判多正确答案问题,尤其在北非、南美和中东表现差。
稳健、多样且具挑战性的文化知识评测基准对衡量大语言模型在多元文化下的表现至关重要。我们提出CulturalBench:一个包含1,696道由人类撰写并人工验证的问题集合,覆盖45个全球区域,包括孟加拉国、津巴布韦、秘鲁等代表性不足地区。每道题经五名独立标注者验证,涵盖从饮食偏好到问候礼仪在内的17个主题。该基准采用受人机红队思想启发的构建方法。与人类表现(92.4%准确率)相比,最难题目对当前顶尖前沿模型仍极具挑战,准确率仅为28.7%至61.5%。我们发现,模型常因多正确答案问题(如“中国人通常用什么餐具?”)而过度拟合单一答案。结果表明,GPT-4o在跨文化任务中显著优于其他模型,超越本地模型(如Mistral在欧洲文化、DeepSeek在中文文化上的表现)。总体而言,模型在北非、南美洲和中东相关问题上表现最弱。
原文摘要 · Abstract (English)
Robust, diverse, and challenging cultural knowledge benchmarks are essential for measuring our progress towards making LMs that are helpful across diverse cultures. We introduce CulturalBench: a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. Questions are each verified by five independent annotators and span 17 diverse topics ranging from food preferences to greeting etiquette. We construct CulturalBench using methods inspired by Human-AI Red-Teaming. Compared to human performance (92.4% accuracy), the hard version of CulturalBench is challenging even for the best-performing frontier LMs, ranging from 28.7% to 61.5% in accuracy. We find that LMs often struggle with tricky questions that have multiple correct answers (e.g., What utensils do the Chinese usually use?), revealing a tendency to overfit to a single answer. Our results indicate that GPT-4o substantially outperform other models across cultures, besting local providers (e.g., Mistral on European culture and DeepSeek on Chinese culture). Across the board, models under-perform on questions related to North Africa, South America and Middle East.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。