arXiv:2506.07661cs.LGcs.IT2025-06被引 2

用信息论解释现代机器学习为何有效,关键在模型复杂度的广度。

Information-Theoretic Framework for Understanding Modern Machine-Learning

  • 以对数损失下的通用预测为视角,用模型邻域体积定义架构复杂度。
  • 复杂度范围广的模型能适应高度过参数化的学习场景。
  • 适合研究深度学习机制、优化原理及新架构设计的人。

我们提出一种信息论框架,将学习视为对数损失下的通用预测,通过后悔界刻画。核心是基于架构的模型复杂度概念,定义为数据生成过程附近或其在模型类上的投影的模型概率质量或体积。该体积与期望海森矩阵或费舍尔信息矩阵的谱特性相关,可进行可计算近似。我们认为成功架构具有宽广的复杂度范围,使其能在高度过参数化模型类中学习。该框架揭示了归纳偏置的作用、随机梯度下降的有效性以及平坦极小值等现象。它统一了在线、批量、监督和生成设置,适用于随机可实现与非可实现两种情形。此外,它为深度神经网络和Transformer等现代架构的成功提供了洞见,指出其分层结构自然带来宽广的复杂度范围。这些见解为设计性能相当甚至更优的替代架构开辟了道路。

原文摘要 · Abstract (English)

We introduce an information-theoretic framework that views learning as universal prediction under log loss, characterized through regret bounds. Central to the framework is an effective notion of architecture-based model complexity, defined by the probability mass or volume of models in the vicinity of the data-generating process, or its projection on the model class. This volume is related to spectral properties of the expected Hessian or the Fisher Information Matrix, leading to tractable approximations. We argue that successful architectures possess a broad complexity range, enabling learning in highly over-parameterized model classes. The framework sheds light on the role of inductive biases, the effectiveness of stochastic gradient descent, and phenomena such as flat minima. It unifies online, batch, supervised, and generative settings, and applies across the stochastic-realizable and agnostic regimes. Moreover, it provides insights into the success of modern machine-learning architectures, such as deep neural networks and transformers, suggesting that their broad complexity range naturally arises from their layered structure. These insights open the door to the design of alternative architectures with potentially comparable or even superior performance.

信息论模型复杂度深度学习优化机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。