用布鲁姆分类法解析大模型认知复杂度,发现思维层级可线性区分。
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
- 用布鲁姆分类法分层探测模型激活向量,检验不同认知层级是否可线性分离。
- 在所有认知层级上平均准确率达95%,表明认知复杂度编码于线性可解子空间。
- 适合对模型内部机制、认知建模感兴趣的读者,尤其关注可解释性研究者。
大语言模型的黑箱特性亟需超越表面性能指标的新评估框架。本研究以布鲁姆分类法为层次化视角,分析不同大模型中高维激活向量的内部神经表征,探究从基础记忆(Remember)到抽象综合(Create)等不同认知层级是否在模型残差流中线性可分。结果表明,线性分类器在所有布鲁姆层级上平均准确率达到约95%,强烈支持认知层级编码于模型表征的线性可访问子空间。该发现表明,模型在前向传播早期即已解析提示的认知难度,且各层表征的可分离性随深度递增。
原文摘要 · Abstract (English)
The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. This study investigates the internal neural representations of cognitive complexity using Bloom's Taxonomy as a hierarchical lens. By analyzing high-dimensional activation vectors from different LLMs, we probe whether different cognitive levels, ranging from basic recall (Remember) to abstract synthesis (Create), are linearly separable within the model's residual streams. Our results demonstrate that linear classifiers achieve approximately 95% mean accuracy across all Bloom levels, providing strong evidence that cognitive level is encoded in a linearly accessible subspace of the model's representations. These findings provide evidence that the model resolves the cognitive difficulty of a prompt early in the forward pass, with representations becoming increasingly separable across layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。