arXiv:2607.13432cs.LG2026-07

提出信息论度量局部冗余,更好预测模型持续学习能力。

Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization

论文配图:Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization
图 1 · 摘自论文原文
  • 基于压缩理论定义局部冗余,衡量模型微调潜力。
  • 合成记忆任务中梯度平方期望可作为下界,高效计算。
  • 在图像与时序迁移任务中优于传统指标,助选最佳预训练节点。

可塑性——神经网络适应新任务的能力——对持续学习和迁移学习至关重要。现有度量如有效秩、死神经元比例和权重范数缺乏理论基础,且与新任务表现相关性差。本文提出局部冗余,一种源自通用压缩理论的信息论度量。局部冗余定义为沿梯度方向的无穷小邻域内局部模型族的最坏情况冗余,具有可塑性的理论依据。尽管局部冗余难以精确计算,我们证明合成记忆任务上的期望平方梯度范数可提供高效可计算的下界。在持续图像分类和时间序列迁移学习实验中,局部冗余比现有度量更准确预测下游性能,并可用于验证损失趋于平稳时的预训练检查点选择。

原文摘要 · Abstract (English)

Plasticity -- a neural network's ability to adapt to new tasks -- is critical for continual and transfer learning. Existing measures, such as effective rank, dead neuron fraction, and weight norm, lack theoretical grounding and correlate poorly with performance on new tasks. We introduce local redundancy, an information-theoretic measure derived from universal compression theory. We define local redundancy as the worst-case redundancy of a local model family -- parameters in an infinitesimal neighborhood along gradient directions -- and show this is a principled measure of plasticity. Although local redundancy is intractable to compute exactly, we prove that the expected squared gradient norm on a synthetic memorization task provides an efficiently computable lower bound. Experiments on continual image classification and time series transfer learning demonstrate that local redundancy predicts downstream performance better than existing measures and enables pretraining checkpoint selection where validation loss plateaus.

可塑性信息论持续学习模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。