用概念树解析大模型如何形成和稳定概念
MindCraft: How Concept Trees Take Shape In Deep Models
- 通过谱分解构建概念路径,追踪概念在层间的层级演化
- 在医疗诊断等多领域验证,能准确恢复语义层级与解耦潜在概念
- 为理解大模型内部表征提供可解释的新框架,适合研究者使用
大规模基础模型在语言、视觉和推理任务中表现优异,但其内部如何结构化并稳定概念仍不明确。受因果推断启发,我们提出基于概念树的MindCraft框架。通过在每一层应用谱分解,并将主方向连接成分支概念路径,概念树重构了概念从共享表示中分化为线性可分子空间的层级演化过程,精确揭示了概念分离的时间点。在医学诊断、物理推理和政治决策等多个跨学科场景的实证评估表明,概念树能够恢复语义层级、解耦潜在概念,并在多个领域具有广泛适用性。该框架为深入分析深度模型中的概念表征提供了通用而强大的工具,标志着可解释人工智能基础的重要进展。
原文摘要 · Abstract (English)
Large-scale foundation models demonstrate strong performance across language, vision, and reasoning tasks. However, how they internally structure and stabilize concepts remains elusive. Inspired by causal inference, we introduce the MindCraft framework built upon Concept Trees. By applying spectral decomposition at each layer and linking principal directions into branching Concept Paths, Concept Trees reconstruct the hierarchical emergence of concepts, revealing exactly when they diverge from shared representations into linearly separable subspaces. Empirical evaluations across diverse scenarios across disciplines, including medical diagnosis, physics reasoning, and political decision-making, show that Concept Trees recover semantic hierarchies, disentangle latent concepts, and can be widely applied across multiple domains. The Concept Tree establishes a widely applicable and powerful framework that enables in-depth analysis of conceptual representations in deep models, marking a significant step forward in the foundation of interpretable AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。