arXiv:2605.09294cs.LGcs.AI2026-05

用宏观状态描述大模型计算,让推理过程可解释可干预

Towards Effective Theory of LLMs: A Representation Learning Approach

论文配图:Towards Effective Theory of LLMs: A Representation Learning Approach
图 1 · 摘自论文原文
  • 通过自监督学习从隐藏层轨迹中提取高阶宏观变量
  • 宏观变量能捕捉语义结构并提前预测如奉承等行为
  • 可追踪推理心理状态,为可控生成提供因果控制手段

我们提出表征有效理论(RET),一种以学习到的宏观状态而非微观细节来描述大语言模型计算的框架。RET利用类似BYOL/JEPA的自监督目标,从隐藏状态轨迹中学习这些宏观状态,将激活值粗粒化为保留预测与解释相关高层结构的宏观变量。我们评估这些宏观变量在可解释性上的实际意义:RET生成时间上一致的状态,揭示了推理的“心智状态”轨迹,捕捉高层语义结构,支持对奉承等行为结果的早期预测,并为引导生成进入可解释的计算阶段提供因果控制。这些结果表明,通过RET可以建立对大模型计算的有效描述:即具有动态意义的高层次变量,支持解释、预测与干预。

原文摘要 · Abstract (English)

We propose Representational Effective Theory (RET), a framework for describing large language model computation in terms of learned macrostates rather than microscopic details. RET learns these macrostates from hidden-state trajectories using a BYOL/JEPA-style self-supervised objective, coarse-graining activations into macrovariables that preserve higher-level structure relevant for prediction and interpretation. We evaluate whether these macrovariables are practically relevant for interpretability: RET yields temporally consistent states that reveal "mental-state" trajectories of reasoning, capture high-level semantic structure, support early prediction of behavioral outcomes such as sycophancy, and provide causal handles for steering generations toward interpretable computational phases. Together, these results suggest that LLM computation admits useful effective descriptions via RET: high-level, dynamically meaningful variables that support interpretation, prediction, and intervention.

大模型解释表示学习宏观变量可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。