arXiv:2510.23006cs.CLcs.AI2025-10

对比不同模型架构的上下文学习机制,发现注意力与Mamba层是关键。

Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architectures

  • 通过行为探测与干预分析,研究Transformer、状态空间与混合模型的上下文学习
  • 功能向量主要位于自注意力和Mamba层,对参数化知识检索更重要
  • 揭示了不同架构在内部机制上的差异,适合关注模型可解释性的研究者

我们在两类基于知识的上下文学习任务上,对最先进的Transformer、状态空间及混合大语言模型进行了深入评估。结合行为探测与干预方法,发现尽管不同架构的模型在任务表现上可能相似,其内部机制仍存在差异。我们发现负责上下文学习的功能向量(FVs)主要集中在自注意力层和Mamba层;推测Mamba2可能采用不同于FVs的机制实现上下文学习。FVs对涉及参数化知识检索的任务更为重要,但对上下文知识理解作用较小。本研究促进了对不同架构与任务类型下上下文学习机制的更细致理解。方法上,强调结合行为分析与机制分析的重要性,以全面探究大模型能力。

原文摘要 · Abstract (English)

We perform in-depth evaluations of in-context learning (ICL) on state-of-the-art transformer, state-space, and hybrid large language models over two categories of knowledge-based ICL tasks. Using a combination of behavioral probing and intervention-based methods, we have discovered that, while LLMs of different architectures can behave similarly in task performance, their internals could remain different. We discover that function vectors (FVs) responsible for ICL are primarily located in the self-attention and Mamba layers, and speculate that Mamba2 uses a different mechanism from FVs to perform ICL. FVs are more important for ICL involving parametric knowledge retrieval, but not for contextual knowledge understanding. Our work contributes to a more nuanced understanding across architectures and task types. Methodologically, our approach also highlights the importance of combining both behavioural and mechanistic analyses to investigate LLM capabilities.

上下文学习模型机制状态空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。