冗余不是浪费,而是学习系统稳定性的关键结构特征。
Redundancy as a Structural Information Principle for Learning and Generalization
- 将冗余重新定义为信息组织的几何结构,统一多种传统度量
- 发现冗余存在上下界,平衡过压缩与过耦合时泛化最优
- 实验验证掩码自编码器在最佳冗余处泛化能力最强
我们提出一个理论框架,将经典信息论扩展至有限且结构化的系统,将冗余重新定义为信息组织的根本属性而非低效表现。在此框架中,冗余表现为一组通用的信息分歧,统一了互信息、卡方依赖性及谱冗余等多种经典度量,揭示它们实为共享冗余几何的投影。该理论进一步预测冗余具有上下界,形成一个平衡过压缩(结构损失)与过耦合(坍塌)的最优均衡点。尽管经典通信理论偏好最小冗余以提高传输效率,但真实世界的学习系统——这类有限且有结构的系统——在接近此均衡点时表现出最大稳定性与泛化能力。通过掩码自编码器的实验验证了这一原理:模型在稳定冗余水平下达到最佳泛化性能。这些结果确立了冗余作为可测量、可调节的量,连接了通信的渐近世界与学习的有限世界。
原文摘要 · Abstract (English)
We present a theoretical framework that extends classical information theory to finite and structured systems by redefining redundancy as a fundamental property of information organization rather than inefficiency. In this framework, redundancy is expressed as a general family of informational divergences that unifies multiple classical measures, such as mutual information, chi-squared dependence, and spectral redundancy, under a single geometric principle. This reveals that these traditional quantities are not isolated heuristics but projections of a shared redundancy geometry. The theory further predicts that redundancy is bounded both above and below, giving rise to an optimal equilibrium that balances over-compression (loss of structure) and over-coupling (collapse). While classical communication theory favors minimal redundancy for transmission efficiency, finite and structured systems, such as those underlying real-world learning, achieve maximal stability and generalization near this equilibrium. Experiments with masked autoencoders are used to illustrate and verify this principle: the model exhibits a stable redundancy level where generalization peaks. Together, these results establish redundancy as a measurable and tunable quantity that bridges the asymptotic world of communication and the finite world of learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。