用解释评估多语言模型的文化推理能力,发现语言影响文化理解。
CRaFT: An Explanation-Based Framework for Evaluating Cultural Reasoning in Multilingual Language Models
- 基于解释内容设计四维评估指标,超越单纯正确率。
- 跨语言测试显示阿拉伯语降低文化流畅性,孟加拉语提升,西班牙语稳定。
- GPT适应性强但不一致,FANAR稳定但僵化,适合改进文化适配模型。
正确答案未必代表文化理解。我们提出CRaFT,一种基于解释的多语言评估框架,用于检验大语言模型(LLMs)在跨文化情境下的推理能力。不同于仅以准确率评分,CRaFT通过文化流畅性、偏离度、一致性与语言适应性四个可解释指标评估模型解释。我们在世界价值观调查中选取50个与文化相关的题目,翻译成阿拉伯语、孟加拉语和西班牙语,对GPT、DeepSeek和FANAR三个模型在超过2,100个回答-解释对上进行评估。结果表明,跨语言推理存在显著差异:阿拉伯语环境下文化流畅性下降,孟加拉语环境下提升,西班牙语则保持稳定。GPT在跨语言适应上表现更好,但一致性较低;FANAR表现出稳定但僵化的推理模式。研究提示,模型的文化意识并非固有,而是由语言表达方式所塑造。CRaFT为多语言场景下的跨文化推理评估提供了新视角,有助于构建更具文化适应性的语言模型。
原文摘要 · Abstract (English)
Correct answers do not necessarily reflect cultural understanding. We introduce CRaFT, an explanation-based multilingual evaluation framework designed to assess how large language models (LLMs) reason across cultural contexts. Rather than scoring outputs solely based on accuracy, CRaFT evaluates model explanations using four interpretable metrics: Cultural Fluency, Deviation, Consistency, and Linguistic Adaptation. We apply the framework to 50 culturally grounded questions from the World Values Survey, translated into Arabic, Bengali, and Spanish, and evaluate three models (GPT, DeepSeek, and FANAR) across over 2,100 answer-explanation pairs. Results reveal significant cross-lingual variation in reasoning: Arabic reduces fluency, Bengali enhances it, and Spanish remains largely stable. While GPT adapts more effectively across languages, it exhibits lower consistency; FANAR shows stable but rigid reasoning. These findings suggest that cultural awareness in LLMs is not intrinsic but emerges through linguistic framing. CRaFT offers a new lens for evaluating cross-cultural reasoning in multilingual settings, providing actionable insights for building culturally adaptive language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。