arXiv:2509.23323cs.LG2025-09NeurIPS被引 5

提出可识别的时序因果表示框架,提升大模型内部机制可解释性

LLM Interpretability with Identifiable Temporal-Instantaneous Representation

  • 融合时序与瞬时因果关系建模,构建可识别的表示学习框架
  • 在模拟数据上验证有效,成功发现大模型激活中的有意义概念关联
  • 为大模型可解释性提供理论保障,适合研究者深入分析模型内部机制

尽管大语言模型具备卓越能力,但理解其内部表征仍具挑战。机制可解释性工具如稀疏自编码器(SAEs)虽能提取可解释特征,却缺乏时序依赖建模、瞬时关系表达及理论保证,削弱了后续分析的理论基础与实践信心。因果表征学习(CRL)虽具理论根基,但现有方法因计算效率低,难以扩展至大模型丰富的概念空间。为此,我们提出专为高维概念空间设计的可识别时序因果表征学习框架,捕捉时滞与瞬时因果关系。该方法提供理论保障,并在逼近真实世界复杂度的合成数据集上验证有效性。通过将SAE技术与我们的时序因果框架结合,成功在大模型激活中发现有意义的概念关系。结果表明,同时建模时序与瞬时概念关系可显著提升大模型可解释性。

原文摘要 · Abstract (English)

Despite Large Language Models' remarkable capabilities, understanding their internal representations remains challenging. Mechanistic interpretability tools such as sparse autoencoders (SAEs) were developed to extract interpretable features from LLMs but lack temporal dependency modeling, instantaneous relation representation, and more importantly theoretical guarantees, undermining both the theoretical foundations and the practical confidence necessary for subsequent analyses. While causal representation learning (CRL) offers theoretically grounded approaches for uncovering latent concepts, existing methods cannot scale to LLMs' rich conceptual space due to inefficient computation. To bridge the gap, we introduce an identifiable temporal causal representation learning framework specifically designed for LLMs' high-dimensional concept space, capturing both time-delayed and instantaneous causal relations. Our approach provides theoretical guarantees and demonstrates efficacy on synthetic datasets scaled to match real-world complexity. By extending SAE techniques with our temporal causal framework, we successfully discover meaningful concept relationships in LLM activations. Our findings show that modeling both temporal and instantaneous conceptual relationships advances the interpretability of LLMs.

大模型可解释性因果表征时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。