用独立成分分析揭示大模型隐藏状态中的语义特征
Transforming Hidden States into Binary Semantic Features
- 通过独立成分分析解耦隐藏状态中的语义成分
- 发现大模型隐藏状态中确实编码了可解释的语义特征
- 为理解大模型内部语义表示提供新视角
大规模语言模型虽源自分布语义理论,但近年来似乎与其渐行渐远。本文重新引入分布语义理论,利用独立成分分析(ICA)克服其应用挑战,证明大语言模型的隐藏状态中确实蕴含语义特征。研究发现,这些隐藏状态能够被分解为具有明确语义意义的独立成分,表明分布语义理论仍可用于解析现代大模型的内部表示。该方法有助于理解模型如何组织和表征语言知识。
原文摘要 · Abstract (English)
Large language models follow a lineage of many NLP applications that were directly inspired by distributional semantics, but do not seem to be closely related to it anymore. In this paper, we propose to employ the distributional theory of meaning once again. Using Independent Component Analysis to overcome some of its challenging aspects, we show that large language models represent semantic features in their hidden states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。