发现有序回归中的神经坍缩现象,揭示了深层模型的几何规律。
Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model
- 用无约束特征模型分析有序回归,提出有序神经坍缩新机制。
- 证明三类坍缩特性:类内特征收敛、类均值对齐、潜变量有序排列。
- 适合研究深度学习几何性质或优化有序回归的学者参考。
深度分类任务中著名的神经坍缩(NC)现象,表现为倒数第二层特征与最终分类器呈现极简几何结构,近年来引发广泛关注。无约束特征模型(UFM)被提出用于理论解释该现象,且已有大量研究将NC拓展至非分类任务并应用于实际场景。本文结合累积链接模型与UFM,探究深度有序回归(OR)中是否存在类似现象。结果表明,确实存在一种称为有序神经坍缩(ONC)的新现象,具有三个特性:(ONC1) 在正则化下,同一类别的最优特征坍缩至类内均值;(ONC2) 这些类均值与分类器对齐,坍缩至一维子空间;(ONC3) 最优潜在变量(对应分类任务中的logits)按类别顺序对齐,尤其在零正则化极限下,潜在变量与阈值间呈现出高度局部且简单的几何关系。我们在固定阈值条件下,于UFM框架内严格证明上述性质,并在多个数据集上实证验证。进一步讨论这些洞察如何应用于有序回归,强调固定阈值的潜在价值。
原文摘要 · Abstract (English)
A phenomenon known as ''Neural Collapse (NC)'' in deep classification tasks, in which the penultimate-layer features and the final classifiers exhibit an extremely simple geometric structure, has recently attracted considerable attention, with the expectation that it can deepen our understanding of how deep neural networks behave. The Unconstrained Feature Model (UFM) has been proposed to explain NC theoretically, and there emerges a growing body of work that extends NC to tasks other than classification and leverages it for practical applications. In this study, we investigate whether a similar phenomenon arises in deep Ordinal Regression (OR) tasks, via combining the cumulative link model for OR and UFM. We show that a phenomenon we call Ordinal Neural Collapse (ONC) indeed emerges and is characterized by the following three properties: (ONC1) all optimal features in the same class collapse to their within-class mean when regularization is applied; (ONC2) these class means align with the classifier, meaning that they collapse onto a one-dimensional subspace; (ONC3) the optimal latent variables (corresponding to logits or preactivations in classification tasks) are aligned according to the class order, and in particular, in the zero-regularization limit, a highly local and simple geometric relationship emerges between the latent variables and the threshold values. We prove these properties analytically within the UFM framework with fixed threshold values and corroborate them empirically across a variety of datasets. We also discuss how these insights can be leveraged in OR, highlighting the use of fixed thresholds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。