分析大模型如何处理词语组合,发现越大越碎片化
Interpreting token compositionality in LLMs: A robustness analysis
- 通过成分感知池化,干预不同层的激活值以研究组合语义
- 模型越大,词语成分信息越分散,无法形成统一语义表示
- 适合关注模型可解释性与架构缺陷的研究者
理解大语言模型(LLMs)的内部机制对提升其可靠性、可解释性和推理过程至关重要。我们提出成分感知池化(CAP),一种基于组合性、机制可解释性和信息论的方法,通过在模型各层级进行基于成分的池化来系统干预激活值,分析模型如何处理组合性语言结构。在反向定义建模、上下位词与同义词预测任务上的实验揭示了变压器模型在处理组合抽象时的关键局限:没有特定层级能基于成分部分将词汇整合为统一的语义表征。我们观察到信息处理的碎片化现象,且随模型规模增大而加剧,表明更大模型在这些干预中表现更差,信息分散程度更高。这种碎片化可能源于变压器的训练目标与架构设计,导致难以形成系统且连贯的表征。研究结果突显了当前变压器架构在组合语义处理和模型可解释性方面的根本局限,强调需发展新方法以应对这些挑战。
原文摘要 · Abstract (English)
Understanding the internal mechanisms of large language models (LLMs) is integral to enhancing their reliability, interpretability, and inference processes. We present Constituent-Aware Pooling (CAP), a methodology designed to analyse how LLMs process compositional linguistic structures. Grounded in principles of compositionality, mechanistic interpretability, and information theory, CAP systematically intervenes in model activations through constituent-based pooling at various model levels. Our experiments on inverse definition modelling, hypernym and synonym prediction reveal critical insights into transformers' limitations in handling compositional abstractions. No specific layer integrates tokens into unified semantic representations based on their constituent parts. We observe fragmented information processing, which intensifies with model size, suggesting that larger models struggle more with these interventions and exhibit greater information dispersion. This fragmentation likely stems from transformers' training objectives and architectural design, preventing systematic and cohesive representations. Our findings highlight fundamental limitations in current transformer architectures regarding compositional semantics processing and model interpretability, underscoring the critical need for novel approaches in LLM design to address these challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。