用几何胶囊定位模型内部概念,可追踪其动态变化。
Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

- 用可解释的几何胶囊拟合概念在表示空间中的形状。
- 发现预训练与强化学习后概念空间变化模式迥异。
- 适合研究模型可解释性与动态表征演化的研究人员。
理解机器学习模型内部表示中概念的编码方式,是机制可解释性的核心问题,对深度学习科学和可信部署日益强大的模型至关重要。现有方法主要将表示映射到更易解释的空间,但无法直接刻画概念在表示空间中的占据方式;诸多假设缺乏严格验证,且多集中于静态表示。本文提出Capsule Lens框架,将概念占据的区域匹配为由若干可解释参数定义的简单、可追踪的几何形式——胶囊,通过闭式求解拟合每个概念的几何特征,并在保留样本上进行验证。该方法应用于静态与动态表示两大场景:在静态表示中,展示如何定位不同模型中的概念几何,并通过跨度与范数曲线揭示关键几何特性;在动态表示中,通过三个案例研究追踪不同训练设置引发的表示漂移:CLIP预训练、视觉问答任务的强化学习后训练、数学推理任务的强化学习后训练。分析显示,几何动态存在显著差异,从CLIP预训练中的广泛网络重构到强化学习后训练中的局部化、概念特定变化。结果既包含与已有文献一致的发现,也呈现新颖观察。我们认为,Capsule Lens是定位、分析和追踪静态与动态表示中概念几何的有力工具。
原文摘要 · Abstract (English)
Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule, defined by several interpretable parameters, fitted in closed form to each concept's geometry and validated on held-out samples. We apply Capsule Lens in two major settings: static and dynamic representations. On static representations, we demonstrate how to locate concept geometry across various models, and how the span and norm curves uncover important geometric characteristics. On dynamic representations, we present three case studies tracking representation drifts induced by distinct training settings, CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning. These analyses reveal qualitatively different geometric dynamics, ranging from broad network-wide restructuring in CLIP pretraining to localized and concept-specific changes in RL post-training. Our results include findings aligned with existing literature as well as novel observations. We believe Capsule Lens stands as a promising tool for locating, analyzing, and tracking concept geometry in both static and dynamic representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。