用几何方法分析模型结构,揭示神经网络的内在空间特性。
The Geometry of Machine Learning Models
- 将模型划分表示为黎曼单纯复形,捕捉单元体积与夹角等几何特征。
- 通过微分形式追踪各层几何结构,实现可计算的几何正则化。
- 提出离散曲率度量,用于诊断模型复杂性与学习动态。
本文提出一个数学框架,通过分析机器学习模型所诱导划分的几何特性来理解其结构。将划分表示为黎曼单纯复形,不仅刻画邻接关系,还包含单元体积、面体积及相邻单元间的二面角等几何属性。针对神经网络,引入微分形式方法,利用拉回运算在各层间追踪几何结构,聚焦于含数据的单元以保证计算可行性。该框架支持直接惩罚不良空间配置的几何正则化,并提供基于扩展拉普拉斯算子与单纯复形样条的新工具用于模型优化。我们进一步研究数据分布如何在模型划分中诱发有效几何曲率,构建顶点的离散曲率度量以量化局部几何复杂性,以及边上的统计里奇曲率以表征单元对之间的关系。尽管聚焦数学基础,这一几何视角为模型解释、正则化和学习动态诊断提供了新途径。
原文摘要 · Abstract (English)
This paper presents a mathematical framework for analyzing machine learning models through the geometry of their induced partitions. By representing partitions as Riemannian simplicial complexes, we capture not only adjacency relationships but also geometric properties including cell volumes, volumes of faces where cells meet, and dihedral angles between adjacent cells. For neural networks, we introduce a differential forms approach that tracks geometric structure through layers via pullback operations, making computations tractable by focusing on data-containing cells. The framework enables geometric regularization that directly penalizes problematic spatial configurations and provides new tools for model refinement through extended Laplacians and simplicial splines. We also explore how data distribution induces effective geometric curvature in model partitions, developing discrete curvature measures for vertices that quantify local geometric complexity and statistical Ricci curvature for edges that captures pairwise relationships between cells. While focused on mathematical foundations, this geometric perspective offers new approaches to model interpretation, regularization, and diagnostic tools for understanding learning dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。