arXiv:2505.24635cs.CL2025-05ACL被引 18

拆解语言与文化维度,发现多语言模型表现受文化语境影响

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

  • 构建语言与文化双维度评估框架,分离跨语言与跨文化测试
  • 发现语言与文化匹配时模型性能更优,存在‘文化-语言协同’现象
  • 特定神经元激活比例可作训练期多语言能力评估指标,适合模型开发者

本文提出一种双维度评估框架,全面评测大语言模型的多语言能力。通过将评估分解为语言媒介和文化语境两个维度,该框架能细致分析模型在母语及跨文化语境下的问答表现。在多种模型上进行广泛评估后,发现显著的‘文化-语言协同’现象:当问题与语言的文化背景一致时,模型表现更优。通过可解释性探测进一步发现,在语言对应的文化语境中,更多特定神经元被激活。该激活比例或可作为模型训练期间多语言能力的潜在评估指标。研究挑战了大语言模型主要基于英语数据训练、各语言表现均一的普遍认知,强调文化和语言双重评估的必要性。代码见 https://yingjiahao14.github.io/Dual-Evaluation/

原文摘要 · Abstract (English)

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framework enables a nuanced analysis of LLMs' ability to process questions within both native and cross-cultural contexts cross-lingually. Extensive evaluations are conducted on a wide range of models, revealing a notable "CulturalLinguistic Synergy" phenomenon, where models exhibit better performance when questions are culturally aligned with the language. This phenomenon is further explored through interpretability probing, which shows that a higher proportion of specific neurons are activated in a language's cultural context. This activation proportion could serve as a potential indicator for evaluating multilingual performance during model training. Our findings challenge the prevailing notion that LLMs, primarily trained on English data, perform uniformly across languages and highlight the necessity of culturally and linguistically model evaluations. Our code can be found at https://yingjiahao14. github.io/Dual-Evaluation/.

多语言模型文化语境评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。