构建SAGE基准测试,评估大模型跨文化理解真实水平。
Do Large Language Models Truly Understand Cross-cultural Differences?
- 基于文化理论设计九维框架,对210个核心概念进行跨文化对齐
- 创建4530个测试题,覆盖15个真实场景,揭示模型系统性短板
- 支持多语言扩展,适合研究跨文化认知与AI伦理的学者使用
近年来,大语言模型在多语言任务中表现强劲,跨文化理解能力成为关键竞争力。然而,现有评估基准存在三方面局限:缺乏情境化场景、跨文化概念映射不足、深层文化推理能力有限。为此,我们提出SAGE——一个基于情景的基准,通过跨文化核心概念对齐与生成式任务设计,评估模型的跨文化理解与推理能力。基于文化理论,将跨文化能力划分为九个维度,整理出210个核心概念,在15个具体现实场景中构建了4530个测试项,归入四类跨文化情境,遵循标准题目设计原则。SAGE支持持续扩展,实验验证其可迁移至其他语言。结果揭示模型在多个维度和场景下存在明显缺陷,暴露出系统性跨文化推理局限。尽管取得进展,大模型仍远未达到真正细腻的跨文化理解水平。为遵守匿名政策,数据与代码已放入补充材料,未来版本将公开上线。
原文摘要 · Abstract (English)
In recent years, large language models (LLMs) have demonstrated strong performance on multilingual tasks. Given its wide range of applications, cross-cultural understanding capability is a crucial competency. However, existing benchmarks for evaluating whether LLMs genuinely possess this capability suffer from three key limitations: a lack of contextual scenarios, insufficient cross-cultural concept mapping, and limited deep cultural reasoning capabilities. To address these gaps, we propose SAGE, a scenario-based benchmark built via cross-cultural core concept alignment and generative task design, to evaluate LLMs' cross-cultural understanding and reasoning. Grounded in cultural theory, we categorize cross-cultural capabilities into nine dimensions. Using this framework, we curated 210 core concepts and constructed 4530 test items across 15 specific real-world scenarios, organized under four broader categories of cross-cultural situations, following established item design principles. The SAGE dataset supports continuous expansion, and experiments confirm its transferability to other languages. It reveals model weaknesses across both dimensions and scenarios, exposing systematic limitations in cross-cultural reasoning. While progress has been made, LLMs are still some distance away from reaching a truly nuanced cross-cultural understanding. In compliance with the anonymity policy, we include data and code in the supplement materials. In future versions, we will make them publicly available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。