arXiv:2506.16289stat.MLcs.LG2025-06

条件数可衡量神经单元的信息编码效率,指导模型微调防遗忘。

The Condition Number as a Scale-Invariant Proxy for Information Encoding in Neural Units

  • 用条件数衡量神经元对信息的压缩与放大能力。
  • 高条件数对应低信息传递量,体现高效专业化编码。
  • 无需预训练数据即可选择性微调,适合资源受限场景。

本文研究神经网络权重重构张量的条件数与信息编码程度的关系,基于信息论视角提出:高条件数虽非有效编码的充分条件,但可能表明单元已学会选择性放大和压缩信息。针对高斯输入的线性单元,该理论将条件数与变换的对数体积缩放因子关联至输出熵和学习变换的几何特性。分析显示,在固定权重范数下,奇异值分布越集中(条件数越高),整体信息传输越少,体现专业化且高效的编码策略。此外,线性阶段熵界为收缩型逐元素非线性激活提供了后激活信息的上限,支持条件数作为实际神经网络中编码容量的尺度不变代理。实验以名为KappaTune的方法应用于大语言模型在新任务和新模态下的选择性微调,结果表明其有效缓解灾难性遗忘。相比依赖预训练统计的现有方法,该方法无需访问原始数据,突破了常见限制。

原文摘要 · Abstract (English)

This paper explores the relationship between the condition number of a neural network's weight tensor and the extent of information encoded by the associated processing unit, viewed through the lens of information theory. It argues that a high condition number, though not sufficient for effective knowledge encoding, may indicate that the unit has learned to selectively amplify and compress information. This intuition is formalized for linear units with Gaussian inputs, linking the condition number and the transformation's log-volume scaling factor to the characteristics of the output entropy and the geometric properties of the learned transformation. The analysis demonstrates that for a fixed weight norm, a concentrated distribution of singular values (high condition number) corresponds to reduced overall information transfer, indicating a specialized and efficient encoding strategy. Furthermore, the linear stage entropy bound provides an upper limit on post-activation information for contractive, element-wise nonlinearities, supporting the condition number as a scale-invariant proxy for encoding capacity in practical neural networks. An empirical case study applies these principles to guide selective fine-tuning of Large Language Models for both a new task and a new input modality. The experiments show that the proposed method, named KappaTune, effectively mitigates catastrophic forgetting. Unlike many existing catastrophic forgetting mitigation methods that rely on access to pre-training statistics, which are often unavailable, this selective fine-tuning approach offers a way to bypass this common requirement.

条件数信息编码微调大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。