用稀疏自编码器拆解大模型,让黑箱变透明并提升推理能力
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
- 通过字典学习与稀疏自编码器分解大模型神经元
- 实现单义特征提取,识别内部误解并自动优化提示
- 在数学推理和隐喻检测任务中显著提效,适合模型可解释性研究者
大型语言模型(LLMs)传统上被视为黑箱算法,降低了可信度,并阻碍了下游任务性能的提升。本文采用基于字典学习的稀疏自编码器方法,有效分解大模型,从中提取单义特征,替代原本多义的神经元表征。值得注意的是,该方法能识别模型内部的理解偏差,并自动为提示添加注释以改善模型理解。此外,该方法在下游任务中表现出显著性能提升,例如在数学推理和隐喻检测任务中效果增强。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are traditionally viewed as black-box algorithms, therefore reducing trustworthiness and obscuring potential approaches to increasing performance on downstream tasks. In this work, we apply an effective LLM decomposition method using a dictionary-learning approach with sparse autoencoders. This helps extract monosemantic features from polysemantic LLM neurons. Remarkably, our work identifies model-internal misunderstanding, allowing the automatic reformulation of the prompts with additional annotations to improve the interpretation by LLMs. Moreover, this approach demonstrates a significant performance improvement in downstream tasks, such as mathematical reasoning and metaphor detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。