用几何流驱动深度学习,让模型自动压缩表示并提升泛化能力。
A Geometrically-Grounded Drive for MDL-Based Optimization in Deep Learning
- 基于最小描述长度原理构建动态优化驱动力,融合几何流与任务梯度。
- 理论证明描述长度单调下降,支持拓扑变换与指数收敛,计算复杂度为O(N log N)。
- 适合追求模型可解释性、泛化性和自动化简化架构的研究者。
本文提出一种将最小描述长度(MDL)原则融入深度神经网络训练动态的新框架。不再仅作为模型选择标准,我们将其重构为优化过程中的主动驱动因素。核心是基于几何的认知流形,其演化由耦合里奇流支配,并引入从第一性原理推导出的新型MDL驱动项。该驱动受任务损失梯度调制,在数据保真与模型简化间实现无缝协同,训练中主动压缩内部表示。理论证明描述长度单调递减(定理~\ref{thm:convergence}),通过几何手术协议实现有限次拓扑相变(定理~\ref{thm:surgery}, \ref{thm:ultimate_fate}),并出现普适临界行为(定理~\ref{thm:universality})。进一步提供计算高效的算法,每轮复杂度为O(N log N)(定理~\ref{thm:complexity}),具备数值稳定性保障(定理~\ref{thm:stability})和凸性假设下的指数收敛性(定理~\ref{thm:convergence_rate})。在合成回归与分类任务上的实证验证了理论预测,展示了算法在鲁棒泛化与自主模型简化方面的有效性。本工作为实现更自主、泛化性强且可解释的AI系统提供了信息论与几何深度学习融合的坚实路径。
原文摘要 · Abstract (English)
This paper introduces a novel optimization framework that fundamentally integrates the Minimum Description Length (MDL) principle into the training dynamics of deep neural networks. Moving beyond its conventional role as a model selection criterion, we reformulate MDL as an active, adaptive driving force within the optimization process itself. The core of our method is a geometrically-grounded cognitive manifold whose evolution is governed by a \textit{coupled Ricci flow}, enriched with a novel \textit{MDL Drive} term derived from first principles. This drive, modulated by the task-loss gradient, creates a seamless harmony between data fidelity and model simplification, actively compressing the internal representation during training. We establish a comprehensive theoretical foundation, proving key properties including the monotonic decrease of description length (Theorem~\ref{thm:convergence}), a finite number of topological phase transitions via a geometric surgery protocol (Theorems~\ref{thm:surgery}, \ref{thm:ultimate_fate}), and the emergence of universal critical behavior (Theorem~\ref{thm:universality}). Furthermore, we provide a practical, computationally efficient algorithm with $O(N \log N)$ per-iteration complexity (Theorem~\ref{thm:complexity}), alongside guarantees for numerical stability (Theorem~\ref{thm:stability}) and exponential convergence under convexity assumptions (Theorem~\ref{thm:convergence_rate}). Empirical validation on synthetic regression and classification tasks confirms the theoretical predictions, demonstrating the algorithm's efficacy in achieving robust generalization and autonomous model simplification. This work provides a principled path toward more autonomous, generalizable, and interpretable AI systems by unifying geometric deep learning with information-theoretic principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。