arXiv:2510.16565cs.CLcs.AI2025-10中稿 · CIKM 2025 Workshop…

通过激活路径分析,揭示多语言大模型的文化理解机制差异。

Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models

  • 用激活路径重叠度衡量文化理解,对比跨语言同国与同语言异国问题。
  • 同语言异国问题内部路径重叠更高,显示语言模式主导响应差异。
  • 韩朝案例显示语言相似不等于文化认知一致,适合研究跨文化模型偏差。

大语言模型在多元文化场景中广泛应用,准确的文化理解至关重要。以往评估多聚焦输出层面,忽略了驱动回答差异的内在因素;而基于电路分析的研究覆盖语言有限,且极少关注文化维度。本文通过测量在两种条件下回答语义等价问题时的激活路径重叠:固定问题语言、变换目标国家;固定国家、变换问题语言。同时采用同语言不同国家对(如韩朝)以分离语言与文化影响。结果表明,同语言异国问题的内部路径重叠度高于跨语言同国问题,说明语言特征主导模型内部表示。值得注意的是,韩国-朝鲜这对虽语言相近,但路径重叠低且变异性高,表明语言相似性并不保证内部表征一致。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used across diverse cultural contexts, making accurate cultural understanding essential. Prior evaluations have mostly focused on output-level performance, obscuring the factors that drive differences in responses, while studies using circuit analysis have covered few languages and rarely focused on culture. In this work, we trace LLMs' internal cultural understanding mechanisms by measuring activation path overlaps when answering semantically equivalent questions under two conditions: varying the target country while fixing the question language, and varying the question language while fixing the country. We also use same-language country pairs to disentangle language from cultural aspects. Results show that internal paths overlap more for same-language, cross-country questions than for cross-language, same-country questions, indicating strong language-specific patterns. Notably, the South Korea-North Korea pair exhibits low overlap and high variability, showing that linguistic similarity does not guarantee aligned internal representation.

文化理解大模型语言差异激活分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。