arXiv:2607.21774cs.CLcs.AI2026-07

用自编码器探测中文大模型对哥伦比亚身份的隐含认知

Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

论文配图:Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
图 1 · 摘自论文原文
  • 通过自然语言自编码器分析模型层20的激活状态
  • 发现模型对哥伦比亚语境有隐含身份表征,尤其在西班牙语输入中
  • 为低资源语言变体的偏见评估提供新方法,适合关注AI公平性的研究者

大型语言模型可能从细微的语言线索中推断出人口统计属性,即使这些属性未被明确说明。本初步研究考察Qwen2.5-7B-Instruct在处理哥伦比亚西班牙语和英语提示时,是否内含对哥伦比亚身份、社会经济地位或刻板印象信息的表征。我们采用自然语言自编码器(NLA),对每条提示在层20中四个位置四分位的残差流激活进行语义化还原。数据集包含30个提示,分为15组匹配的西班牙语-英语对,涵盖显性哥伦比亚线索、隐性线索及中性对照。研究报告描述性比例与定性证据,而非统计显著效应,重点在于模型输出前是否存在潜在的国籍或刻板印象表征。本工作将激活层级可解释性与对非主流西班牙语变体的偏见评估相连接。

原文摘要 · Abstract (English)

Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts. We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output. This work connects activation-level interpretability with bias evaluation for underrepresented Spanish varieties.

模型可解释性语言偏见哥伦比亚语自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。