发现语言模型中存在文化绑定注意力头,可提升跨文化识别准确率
Cultural Binding Heads in Language Models
- 通过注意力头干预,实现对文化身份与文化元素的精准绑定
- 移除关键注意力头后,绑定强度下降9%-23%,证明其因果作用
- 适合研究文化偏见、模型可解释性及公平性的人士阅读
大语言模型通常对不同文化群体一视同仁,但上下文本应有所区分,这反映出缺乏差异意识。基于Wang等(2025)提出的N4文化挪用基准,我们采用机械可解释性分析和因子设计,在八种模型(四种架构,含base与instruct版本)中识别出每模型2-3个中层注意力头,这些头在文化绑定过程中起因果作用。文化绑定指将文化物品与其对应身份关联的过程。移除这些头上的身份-物品连接后,绑定强度降低9%-23%。这些头可从instruct模型迁移至base模型,表明文化绑定在预训练阶段已形成。α缩放实验显示剂量响应呈梯度变化,当α=2-3时,生成阶段适度放大可使文化差异化准确率提升1-3个百分点,同时保持中性推理基本不变。知识探测任务显示,模型所知是其实际应用的3-5倍,说明瓶颈在于信息路由而非知识本身。
原文摘要 · Abstract (English)
LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (four architectures, base and instruct). Cultural binding is the process of associating cultural items with the appropriate identity. Knockout of the identity-to-item edges on these heads lowers the binding strength by 9-23%. The identified heads transfer from instruct to base models, suggesting that cultural binding is created at pre-training. An $α$-scaling shows a graded dose-response and moderate amplification steering at generation ($α= 2-3$) increases cultural differentiation accuracy by 1-3 pp while leaving neutral reasoning mostly intact. A knowledge probing task shows that models know 3-5 times more than they act upon it, indicating that the bottleneck lies in routing and not knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。