arXiv:2509.07149cs.LGcs.AI2025-09被引 1

提出新方法评估大模型中电路的可信度,判断其是否稳定可靠。

Measuring Uncertainty in Transformer Circuits with Effective Information Consistency

  • 基于局部雅可比与激活计算一致性得分
  • 结合高斯有效信息代理衡量因果涌现程度
  • 无需反向传播,适合快速分析模型内部机制

机制可解释性已识别出大型语言模型中的功能子图,即变压器电路(TCs),这些电路似乎实现了特定算法。然而,我们缺乏一种单次前向传播的正式方法来量化活跃电路何时表现一致且可信。本文借鉴系统理论中的层化/上同调与因果涌现视角,将之专门应用于变压器电路,提出了有效信息一致性评分(EICS)。EICS结合了(i)由局部雅可比和激活计算得到的归一化层化不一致性,以及(ii)从同一前向状态推导出的高斯有效信息代理,用于衡量电路级因果涌现。该方法为白盒、单次前向传播设计,明确表达单元,使得分无量纲。同时提供得分解读指南、计算开销说明(含快速与精确模式)及一个简化验证分析。在大模型任务上的实证验证暂未开展。

原文摘要 · Abstract (English)

Mechanistic interpretability has identified functional subgraphs within large language models (LLMs), known as Transformer Circuits (TCs), that appear to implement specific algorithms. Yet we lack a formal, single-pass way to quantify when an active circuit is behaving coherently and thus likely trustworthy. Building on prior systems-theoretic proposals, we specialize a sheaf/cohomology and causal emergence perspective to TCs and introduce the Effective-Information Consistency Score (EICS). EICS combines (i) a normalized sheaf inconsistency computed from local Jacobians and activations, with (ii) a Gaussian EI proxy for circuit-level causal emergence derived from the same forward state. The construction is white-box, single-pass, and makes units explicit so that the score is dimensionless. We further provide practical guidance on score interpretation, computational overhead (with fast and exact modes), and a toy sanity-check analysis. Empirical validation on LLM tasks is deferred.

可解释性大模型电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。