用二值网络验证学习即压缩,揭示模型训练中的算法简化规律。
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
- 以二值神经网络为代理,用块分解法衡量算法复杂度。
- 训练过程中算法复杂度下降速度比熵更贴近损失变化趋势。
- 为复杂度感知学习提供基于信息论的理论框架,适合研究模型泛化者。
理解与控制神经网络的信息复杂度是机器学习的核心挑战,涉及泛化、优化和模型容量。现有方法多依赖熵损失和统计指标,但难以捕捉网络结构中深层的因果性算法规律。本文转向算法信息论,以二值神经网络(BNNs)作为初步代理,基于算法概率(AP)及其定义的通用分布,从因果角度刻画学习动态。采用可扩展的算法复杂度近似方法——块分解法(BDM),结果表明其对训练过程中的结构变化追踪能力优于熵,在不同模型规模和随机训练运行中均与训练损失保持更强相关性。这些发现支持将训练视为算法压缩过程的观点,即学习对应于结构规律的逐步内化。本工作为学习进程提供了原则性估计,并提出一种基于信息论、复杂性与可计算性的复杂度感知学习与正则化框架。
原文摘要 · Abstract (English)
Understanding and controlling the informational complexity of neural networks is a central challenge in machine learning, with implications for generalization, optimization, and model capacity. While most approaches rely on entropy-based loss functions and statistical metrics, these measures often fail to capture deeper, causally relevant algorithmic regularities embedded in network structure. We propose a shift toward algorithmic information theory, using Binarized Neural Networks (BNNs) as a first proxy. Grounded in algorithmic probability (AP) and the universal distribution it defines, our approach characterizes learning dynamics through a formal, causally grounded lens. We apply the Block Decomposition Method (BDM) -- a scalable approximation of algorithmic complexity based on AP -- and demonstrate that it more closely tracks structural changes during training than entropy, consistently exhibiting stronger correlations with training loss across varying model sizes and randomized training runs. These results support the view of training as a process of algorithmic compression, where learning corresponds to the progressive internalization of structured regularities. In doing so, our work offers a principled estimate of learning progression and suggests a framework for complexity-aware learning and regularization, grounded in first principles from information theory, complexity, and computability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。