arXiv:2508.08879cs.CLcs.AI2025-08被引 4

揭示大模型内部文化推理的演变过程与偏差

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

  • 通过机制可解释性方法探测模型参数中的文化表征
  • 发现文化推理存在三阶段:领域识别→文化定位→答案生成
  • 模型对小众文化反应迟缓且易混淆,反映知识不平衡

大型语言模型在多元文化场景中广泛应用,亟需理解其内部对不同文化的隐式表征。以往研究仅分析输出结果,忽视了模型参数中文化知识的存储与处理机制,无法解释错误回答的成因。为此,我们提出CulTrace——一种基于机制可解释性的方法,用于探查大模型内部的文化知识表示。通过该方法,我们发现文化推理存在一致的三阶段轨迹:先聚焦问题领域,再识别相关文化,最后确定答案。此外,我们证实模型的文化推理存在不平衡现象,表现为对相关文化识别延迟,且对代表性不足的文化更易混淆。

原文摘要 · Abstract (English)

The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awareness in LLMs by analysing their outputs. This approach overlooks how cultures are represented within the model parameters, missing why models generate incorrect responses. To bridge this gap, we propose CulTrace, a mechanistic interpretability-based method that probes the internal representations of LLMs for cultural knowledge. With CulTrace, we inspect how cultural knowledge is processed across layers and how it is integrated during cultural QA. We find a consistent staged trajectory of cultural reasoning. Models first engage with the question's domain, then resolve the relevant culture, and finally narrow in on an answer. We also demonstrate that models' cultural reasoning is imbalanced, showing delayed relevant culture resolution and more confusion with less-represented cultures.

文化推理可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。