arXiv:2609.08322cs.CLcs.AI2026-09

揭示多语言大模型中刻板印象的内部表征与输出机制差异

Tracing Stereotypes from Representation to Output in Multilingual LLMs

  • 通过多种方法分析模型各层特征,定位刻板印象信息分布
  • 3个模型中特征探测峰值早于归因,深度差达36%-53%
  • 仅6%-18%特征具跨语言影响,无完全跨类别通用特征

多语言大模型在不同语言中表现出不同的刻板印象行为,但行为评分无法揭示相关信息在模型中的具体表征位置及其对输出的影响。为探究内部机制,我们对比了线性探测、归因修补、稀疏自编码器(SAEs)和特征消融在 Llama-3.1-8B、Qwen3-8B 与 Gemma-2-9B 中的表现。所有模型中,探测性能峰值显著早于归因,深度差距达36%-53%。保留的 Llama-Scope 特征常与所选社会类别一致,并形成重复的语义家族,但其词汇对齐性和消融效果在不同 SAE 套件间存在差异。在评估的残差流特征中,仅6%-18%符合跨语言无关的标准,且无任何特征具有跨类别无关性。跨语言特征在 Llama-Scope 中平均消融效应更大,但该模式在其他 SAE 套件中不复现。因此,解码性、输出影响及跨语言消融效应需分别测量。

原文摘要 · Abstract (English)

Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all three models, with a separation of 36-53% of model depth. Retained Llama-Scope features often match the social category on which they were selected and form recurring semantic families, but their lexical alignment and ablation effects vary across SAE suites. Only 6-18% of evaluated residual-stream features have language-agnostic effects under our criterion, and none are category-agnostic. Language-agnostic features have larger mean ablation effects in Llama-Scope, but this pattern does not repeat in the other SAE suites. Decodability, output influence, and cross-lingual ablation effects therefore need to be measured separately.

多语言模型刻板印象特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。