剖析大模型如何内化种族线索,揭示偏见背后的神经机制。
Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity Cues
- 通过探针与干预分析模型内部对种族线索的分布式表征
- 发现不同模型对种族相关语义的敏感度差异显著
- 适合关注模型偏见机制与可解释性研究的读者
大语言模型在高风险场景中越来越多地处理涉及种族和民族属性的文本信息,这些属性可能明确提及或通过语境暗示。然而,现有研究多聚焦结果层面的偏差,缺乏对内部机制的深入理解。本文基于两个公开数据集(涵盖毒性生成与临床叙事理解任务),采用可复现的可解释性流程,分析三种开源模型,结合探针测试、神经元层级归因与定向干预。结果显示,对种族线索的敏感性分布在多个内部单元中,且在不同模型间差异明显;这些单元常与显式群体标签、地理、语言、文化及刻板印象关联交织。对特定单元进行干预可改变部分偏差模式,但仍有显著残留效应,表明有效缓解需理解分布式、任务相关的机制,而非仅操作少数神经元。代码已开源:https://github.com/LARK-NLP-Lab/LLM-Bias-Interpretability。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly operate in high-stakes settings where demographic attributes such as race and ethnicity may be explicitly stated or implicitly suggested through textual cues. However, existing studies primarily document outcome-level disparities, offering limited insight into internal mechanisms underlying these effects. We present a mechanistic study of how race and ethnicity cues are represented and operationalized within LLMs. Using two publicly available datasets spanning toxicity-related generation and clinical narrative understanding tasks, we analyze three open-source models with a reproducible interpretability pipeline combining probing, neuron-level attribution, and targeted intervention. We find that sensitivity to demographic cues is distributed across internal units and varies substantially across models. These units often align with entangled semantic facets, including explicit group labels, geography, language, culture, and associations related to stereotypes. Interventions on selected units can change some biased prediction patterns, but substantial residual effects remain, suggesting that effective mitigation requires understanding distributed, task-specific mechanisms rather than manipulating a small set of identified neurons alone. Code: https://github.com/LARK-NLP-Lab/LLM-Bias-Interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。