用数学展开法分析模型计算路径,无需数据即可解析预测来源。
Jet Expansions of Residual Computation
- 用切线流(jets)展开残差计算图,分解不同路径贡献。
- 发现残差深度存在超指数级路径结构,揭示模型内部机制。
- 适合模型可解释性研究、开发与评估,无需额外数据训练。
我们提出一种基于切线流(jets)的残差计算图展开框架,该方法将截断泰勒级数推广为更通用的算子。该方法能系统地分离不同计算路径对模型输出的贡献。相比蒸馏、探针或早期解码等现有技术,本方法仅依赖模型自身,无需数据、训练或采样。实验表明,该框架可统一解释对数几率透镜(logit lens),揭示递归残差深度中存在(超)指数级路径结构,并支持多种应用:例如利用模型计算中提取的n-gram统计信息构建Transformer语言模型草图,以及索引模型毒性知识层级。该方法实现了无需数据的残差计算分析,适用于模型可解释性、开发与评估。
原文摘要 · Abstract (English)
We introduce a framework for expanding residual computational graphs using jets, operators that generalize truncated Taylor series. Our method provides a systematic approach to disentangle contributions of different computational paths to model predictions. In contrast to existing techniques such as distillation, probing, or early decoding, our expansions rely solely on the model itself and requires no data, training, or sampling from the model. We demonstrate how our framework grounds and subsumes logit lens, reveals a (super-)exponential path structure in the recursive residual depth and opens up several applications. These include sketching a transformer large language model with $n$-gram statistics extracted from its computations, and indexing the models' levels of toxicity knowledge. Our approach enables data-free analysis of residual computation for model interpretability, development, and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。