用稀疏自编码器追踪GPT如何学习并编码社会偏见
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
- 结合GPT与稀疏自编码器,追溯主题在训练中的编码过程
- 发现性别与亲属关系是核心主题,且随网络深度不断交织扩展
- 为审计模型中文化偏见的来源提供可扩展的新方法
大型语言模型(LLMs)在海量非结构化语料上训练,其吸收并再现的社会模式与偏见尚不清晰。现有评估多关注输出或激活值,很少回溯到预训练数据。本文提出一个将LLM与稀疏自编码器(SAEs)结合的分析流程,追踪不同主题在训练过程中的编码方式。以19世纪十位女性作家的37部小说为训练集,主题涵盖性别、婚姻、阶级与道德。通过在各层应用SAEs,并以十一类社会与道德概念进行探针分析,成功将稀疏特征映射至人类可理解的概念。结果揭示出稳定的主题骨架(尤其集中在性别与亲属关系),并显示这些关联随网络深度逐步扩展与纠缠。更广泛地,我们主张该LLM+SAEs框架为审计文化假设如何嵌入模型表征提供了可扩展路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are trained on massive, unstructured corpora, making it unclear which social patterns and biases they absorb and later reproduce. Existing evaluations typically examine outputs or activations, but rarely connect them back to the pre-training data. We introduce a pipeline that couples LLMs with sparse autoencoders (SAEs) to trace how different themes are encoded during training. As a controlled case study, we trained a GPT-style model on 37 nineteenth-century novels by ten female authors, a corpus centered on themes such as gender, marriage, class, and morality. By applying SAEs across layers and probing with eleven social and moral categories, we mapped sparse features to human-interpretable concepts. The analysis revealed stable thematic backbones (most prominently around gender and kinship) and showed how associations expand and entangle with depth. More broadly, we argue that the LLM+SAEs pipeline offers a scalable framework for auditing how cultural assumptions from the data are embedded in model representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。