发现语言模型中语义干扰的通用模式,可跨模型预测行为变化
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
- 通过稀疏自编码器定位模型中语义无关却相互干扰的特征对
- 在小模型中干预后,行为改变可稳定转移到大模型上
- 揭示了模型内部表征存在非直觉的共性结构,适合黑箱控制研究
多义性在语言模型中普遍存在,是解释与行为控制的主要挑战。我们利用稀疏自编码器(SAEs)映射两个小型模型(Pythia-70M 和 GPT-2-Small)的多义拓扑结构,识别出语义无关但模型内存在干扰的 SAE 特征对。在四个干预点(提示、词元、特征、神经元)进行操作,并测量下一词预测分布的变化,发现了暴露系统性弱点的多义结构。关键发现:从小模型中提炼出的反直觉干扰模式能可靠转移至更大的指令微调模型(Llama-3.1-8B/70B-Instruct 和 Gemma-2-9B-Instruct),实现无需访问模型内部即可预测的行为转变。这些结果挑战了多义性仅为随机现象的观点,表明干扰结构具有跨规模和模型家族的泛化能力,暗示内部表征存在收敛的高阶组织,其与直觉关联弱,由潜在规律决定,为黑箱控制及人类与人工认知理论提供了新可能。
原文摘要 · Abstract (English)
Polysemanticity is pervasive in language models and remains a major challenge for interpretation and model behavioral control. Leveraging sparse autoencoders (SAEs), we map the polysemantic topology of two small models (Pythia-70M and GPT-2-Small) to identify SAE feature pairs that are semantically unrelated yet exhibit interference within models. We intervene at four foci (prompt, token, feature, neuron) and measure induced shifts in the next-token prediction distribution, uncovering polysemantic structures that expose a systematic vulnerability in these models. Critically, interventions distilled from counterintuitive interference patterns shared by two small models transfer reliably to larger instruction-tuned models (Llama-3.1-8B/70B-Instruct and Gemma-2-9B-Instruct), yielding predictable behavioral shifts without access to model internals. These findings challenge the view that polysemanticity is purely stochastic, demonstrating instead that interference structures generalize across scale and family. Such generalization suggests a convergent, higher-order organization of internal representations, which is only weakly aligned with intuition and structured by latent regularities, offering new possibilities for both black-box control and theoretical insight into human and artificial cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。