arXiv:2510.15980cs.AI2025-10

用认知负荷追踪模型内部资源分配,揭示推理过程中的关键问题。

Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition

  • 将模型内部资源分配建模为三类认知负荷的时变函数。
  • 可预测错误发生点,提升推理效率15%-30%且保持准确率。
  • 适合研究大模型推理机制或优化推理效率的研究者。

我们提出认知负荷痕迹(CLTs)作为深度模型的中层可解释性框架,灵感来自人类认知中的认知负荷理论。CLTs被定义为量化模型内部资源分配的符号化、时变函数,形式上表示为三元随机过程(IL_t, EL_t, GL_t),对应内在负荷、外在负荷和关联负荷。各分量通过注意力熵、键值缓存未命中率、表征离散度和解码稳定性等可观测代理变量实现。我们提出了符号化表达与可视化方法(负荷曲线、单纯形图),支持对推理动态的可解释分析。在推理与规划基准上的实验表明,CLTs能预测错误发生时机,揭示认知策略,并实现基于负荷引导的干预,使推理效率提升15%-30%,同时保持准确率。

原文摘要 · Abstract (English)

We propose \textbf{Cognitive Load Traces} (CLTs) as a mid-level interpretability framework for deep models, inspired by Cognitive Load Theory in human cognition. CLTs are defined as symbolic, temporally varying functions that quantify model-internal resource allocation. Formally, we represent CLTs as a three-component stochastic process $(\mathrm{IL}_t, \mathrm{EL}_t, \mathrm{GL}_t)$, corresponding to \emph{Intrinsic}, \emph{Extraneous}, and \emph{Germane} load. Each component is instantiated through measurable proxies such as attention entropy, KV-cache miss ratio, representation dispersion, and decoding stability. We propose both symbolic formulations and visualization methods (load curves, simplex diagrams) that enable interpretable analysis of reasoning dynamics. Experiments on reasoning and planning benchmarks show that CLTs predict error-onset, reveal cognitive strategies, and enable load-guided interventions that improve reasoning efficiency by 15-30\% while maintaining accuracy.

可解释性认知负荷推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。