用稀疏自编码器分析大模型如何内化宗教、暴力与地理关联
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
- 通过SAE技术解析模型内部激活模式,探测宗教与暴力概念的关联机制
- 伊斯兰相关特征更常触发暴力语言激活,其他宗教则相对均衡
- 地理关联反映现实宗教分布,揭示模型融合事实与刻板印象的内在逻辑
尽管关于大语言模型偏见的研究日益增多,但多数聚焦于性别和种族,对宗教身份的关注较少。本文探讨了宗教在大语言模型中的内部表征方式及其与暴力和地理概念的交叉关系。利用机械可解释性方法与稀疏自编码器(SAEs),通过Neuronpedia API分析五个模型的潜在特征激活情况,测量宗教与暴力相关提示之间的重叠度,并探究激活上下文中的语义模式。结果显示,所有五种宗教在内部表征上均具有相似的凝聚性,但伊斯兰相关的特征更频繁地与暴力语言特征相关联。相比之下,地理关联主要反映真实世界的宗教人口分布,揭示了模型既嵌入了事实分布,也包含文化刻板印象。这些发现凸显了结构化分析在审计模型输出之外内部表示的重要性,这些表示影响模型行为。
原文摘要 · Abstract (English)
Despite growing research on bias in large language models (LLMs), most work has focused on gender and race, with little attention to religious identity. This paper explores how religion is internally represented in LLMs and how it intersects with concepts of violence and geography. Using mechanistic interpretability and Sparse Autoencoders (SAEs) via the Neuronpedia API, we analyze latent feature activations across five models. We measure overlap between religion- and violence-related prompts and probe semantic patterns in activation contexts. While all five religions show comparable internal cohesion, Islam is more frequently linked to features associated with violent language. In contrast, geographic associations largely reflect real-world religious demographics, revealing how models embed both factual distributions and cultural stereotypes. These findings highlight the value of structural analysis in auditing not just outputs but also internal representations that shape model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。