用拓扑分析可视化高维损失曲面,揭示模型性能与学习动态的新规律。
Visualizing Loss Functions as Topological Landscape Profiles
- 引入拓扑数据法构建高维损失景观新表示
- 表现更好的模型其损失拓扑更简单,性能跃迁处变化剧烈
- 适用于图像分割与科学机器学习,助理解超参影响
在机器学习中,损失函数衡量模型预测与真实值之间的差异。对于神经网络模型,可视化损失随参数变化的形态,可揭示损失景观的局部结构(如平滑性)及模型整体特性(如泛化性能)。尽管已有多种损失景观可视化方法,但多数仅限于一两个方向采样,忽略了极高维空间中的潜在信息。本文提出一种基于拓扑数据分析的新表示方法,实现高维损失景观的可视化。通过该拓扑景观轮廓表征,我们发现模型性能越优,其损失景观的拓扑结构越简单;在低性能向高性能跃迁区域,损失景观形状变化更显著。该方法在图像分割(如UNet)和科学机器学习(如物理信息神经网络)中均有应用,为理解不同超参数空间下的学习动态提供了新视角。
原文摘要 · Abstract (English)
In machine learning, a loss function measures the difference between model predictions and ground-truth (or target) values. For neural network models, visualizing how this loss changes as model parameters are varied can provide insights into the local structure of the so-called loss landscape (e.g., smoothness) as well as global properties of the underlying model (e.g., generalization performance). While various methods for visualizing the loss landscape have been proposed, many approaches limit sampling to just one or two directions, ignoring potentially relevant information in this extremely high-dimensional space. This paper introduces a new representation based on topological data analysis that enables the visualization of higher-dimensional loss landscapes. After describing this new topological landscape profile representation, we show how the shape of loss landscapes can reveal new details about model performance and learning dynamics, highlighting several use cases, including image segmentation (e.g., UNet) and scientific machine learning (e.g., physics-informed neural networks). Through these examples, we provide new insights into how loss landscapes vary across distinct hyperparameter spaces: we find that the topology of the loss landscape is simpler for better-performing models; and we observe greater variation in the shape of loss landscapes near transitions from low to high model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。