arXiv:2601.14063cs.CLcs.AI2026-01被引 2

评测大模型跨文化推理能力,发现其在深层文化理解上存在明显短板。

XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

  • 基于文化可见性三层次设计评测任务,覆盖可观察行为到隐性价值观。
  • 8个主流多语言模型在敏感文化类别上表现显著下降(p<0.005)。
  • 揭示同一语言内存在的区域与宗教偏见,适合跨文化NLP研究者使用。

大语言模型的跨文化能力需要理解并适应不同文化背景下的文化特异性项目(CSIs)。然而,当前评估受限于缺乏高质量的带标注语料库及平行跨文化句子对。本文提出XCR-Bench,一个包含4.1千组平行句和1,098个CSIs的跨文化推理基准,涵盖三个推理任务。该基准结合纽马克的CSI框架与霍尔的文化三角理论,支持从显性行为到隐性社会规范与价值观等多层次文化可见性评估。对8个主流多语言模型的实验表明,先进模型在识别与适应特定类别CSIs方面持续存在弱点,表面记忆与显性文化推理间存在差距。在文化敏感类别及深层文化层面,性能显著下降(p<0.005,8/8模型),且适应质量随目标文化与孟加拉地区变体系统性变化,反映单语环境下仍存在区域与族裔-宗教偏见。我们已公开发布语料库与代码,以推动跨文化自然语言处理研究。

原文摘要 · Abstract (English)

Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cultural contexts. However, progress in evaluating this capability remains limited by the lack of high-quality CSI-annotated corpora with parallel cross-cultural sentence pairs. We introduce XCR-Bench, a Cross(X)-Cultural Reasoning Benchmark containing 4.1k parallel sentences and 1,098 CSIs across three reasoning tasks. XCR-Bench integrates Newmark's CSI framework with Hall's Triad of Culture, enabling evaluation across levels of cultural visibility -- from observable practices to implicit social norms and values. Experiments on eight multilingual LLMs show that state-of-the-art models exhibit consistent weaknesses in identifying and adapting specific categories of CSIs, revealing a gap between surface-level recall and explicit cultural reasoning. Performance declines significantly on culturally sensitive categories and deeper cultural levels (p<0.005, 8/8 models), and adaptation quality varies systematically across target cultures and Bengali regional variants, indicating encoded regional and ethno-religious biases even within a single linguistic setting. We publicly release the corpus and code to support future research on cross-cultural NLP.

跨文化推理大模型评测文化偏见多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。