用自编码器解析大模型内部语义,让人类看懂模型如何思考。
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
- 用稀疏自编码器分离出可理解的特征表示。
- 让大模型自动解释这些特征,验证其能控制模型行为。
- 揭示模型在答题和出错时的内在机制,类比人类认知。
近期大型多模态模型(LMMs)在学术与工业界取得显著进展,但其内部神经表征如何被人类理解仍是个问题。本文提出一个通用框架,用于识别与解释LMMs中的语义特征。首先,采用稀疏自编码器(SAE)将表征解耦为人类可理解的特征;随后,设计自动解释框架,由大模型自身解读SAE中学习到的开放语义特征。该方法以LLaVA-OV-72B分析LLaVA-NeXT-8B模型,结果显示这些特征可有效引导模型行为。结果有助于理解LMMs在特定任务(如EQ测试)中的表现优势,揭示其错误本质,并提出修正策略。研究为理解LMMs内部机制提供了新视角,暗示其与人类认知存在相似性。
原文摘要 · Abstract (English)
Recent advances in Large Multimodal Models (LMMs) lead to significant breakthroughs in both academia and industry. One question that arises is how we, as humans, can understand their internal neural representations. This paper takes an initial step towards addressing this question by presenting a versatile framework to identify and interpret the semantics within LMMs. Specifically, 1) we first apply a Sparse Autoencoder(SAE) to disentangle the representations into human understandable features. 2) We then present an automatic interpretation framework to interpreted the open-semantic features learned in SAE by the LMMs themselves. We employ this framework to analyze the LLaVA-NeXT-8B model using the LLaVA-OV-72B model, demonstrating that these features can effectively steer the model's behavior. Our results contribute to a deeper understanding of why LMMs excel in specific tasks, including EQ tests, and illuminate the nature of their mistakes along with potential strategies for their rectification. These findings offer new insights into the internal mechanisms of LMMs and suggest parallels with the cognitive processes of the human brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。