分析自监督模型的线性区域复杂度,发现其更少但更优。
Complexity of Linear Regions in Self-supervised Deep ReLU Networks

- 用样条摄像法追踪数据附近线性区域分布
- 自监督方法区域数更少,精度相当
- 线性区域几何特征可早期预警表征崩溃
近年来对基于ReLU激活函数网络复杂度的研究日益增多。已有研究关注训练过程中产生的分段线性分区(线性区域)数量的变化,但当前工作多局限于有监督训练模型。自监督学习(SSL)通过损失函数直接优化表示空间,以提升模型在多个下游任务中的性能。本文研究了自监督模型生成的线性区域局部分布。我们利用SplineCam提取靠近数据分布的二维多面体,追踪训练过程中区域的数量、面积、偏心率及边界变化。在标准数据集MNIST和FashionMNIST上,对比了有监督、对比学习与自蒸馏方法。结果表明,自监督方法在达到与有监督模型相当精度时,产生的线性区域显著更少;对比方法随时间快速扩展区域,而自蒸馏方法则趋向于合并邻近区域。最后,我们可在几何空间中早期检测到表示崩溃现象。分析表明,多面体度量可作为表示质量与模型性能的可靠指标。
原文摘要 · Abstract (English)
There has been growing interest in studying the complexity of Rectified Linear Unit (ReLU) based activation networks. Recent work investigates the evolution of the number of piecewise-linear partitions (linear regions) that are formed during training. However, current research is limited to examining the complexity of models trained in a supervised way. Self-Supervised Learning (SSL) differs in that it directly optimises the representation space using a loss function to enhance the model's performance across multiple downstream tasks. This study investigates the local distribution of linear regions produced by SSL models. We demonstrate that the evolution of linear regions correlates with the representation quality by utilising SplineCam to extract two-dimensional polytopes near the data distribution. We track the number, area, eccentricity, and boundaries of regions throughout training. The study compares supervised, contrastive, and self-distillation methods over two standard benchmark datasets, MNIST and FashionMNIST. The analysis of the experimental results shows that self-supervised methods create substantially fewer regions to achieve comparable accuracy to supervised models. Contrastive methods rapidly expand regions over time, whereas self-distillation methods tend to consolidate by merging neighbouring regions. Lastly, we can detect representation collapse early within the geometric space of linear regions. Our analysis suggests that polytopal metrics can serve as reliable indicators of representation quality and model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。