研究大模型如何从少量示例中发现并分解隐藏概念。
Latent Concept Disentanglement in Transformer-based Language Models
- 通过可控任务测试,发现模型能识别离散隐藏概念并逐步组合。
- 在数值参数任务中,模型表示空间存在低维子空间,几何结构反映参数化规律。
- 小模型和大模型均能在上下文学习中有效解耦和利用隐含概念,适合可解释性研究者阅读。
当大型语言模型(LLMs)使用上下文学习(ICL)解决新任务时,必须从示范示例中推断出隐藏概念。这引出了一个关键问题:变压器模型是否以及如何在其计算过程中表示这些潜在结构?本研究通过若干受控任务,采用机制可解释性方法进行探究。首先,在具有潜在离散概念的传递推理任务中,模型成功识别出该概念,并实现逐步的概念组合,拓展了以往对单步推理的研究。其次,在由潜在数值概念参数化的任务中,我们发现模型表示空间中存在低维子空间,其几何结构清晰反映了底层参数化关系。总体而言,我们证明了小型和大型模型都能从少数简略示例中学习并解耦利用上下文中的隐藏概念。
原文摘要 · Abstract (English)
When large language models (LLMs) use in-context learning (ICL) to solve a new task, they must infer latent concepts from demonstration examples. This raises the question of whether and how transformers represent latent structures as part of their computation. Our work experiments with several controlled tasks, studying this question using mechanistic interpretability. First, we show that in transitive reasoning tasks with a latent, discrete concept, the model successfully identifies the latent concept and does step-by-step concept composition. This builds upon prior work that analyzes single-step reasoning. Then, we consider tasks parameterized by a latent numerical concept. We discover low-dimensional subspaces in the model's representation space, where the geometry cleanly reflects the underlying parameterization. Overall, we show that small and large models can indeed disentangle and utilize latent concepts that they learn in-context from a handful of abbreviated demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。