通过变量依赖关系学习特征表示,揭示了损失函数与最大相关性的深层联系。
Dependence Induced Representations
- 基于变量依赖性构造特征表示,理论证明其必要充分条件。
- 多种损失函数(如交叉熵、合页损失)可学习到依赖诱导表示。
- 为深度分类器的神经坍缩现象提供统计解释,适合理论研究者。
我们研究从一对随机变量中学习特征表示的问题,重点关注由它们依赖性所诱导的表示。给出了此类依赖诱导表示的充要条件,并揭示其与Hirschfeld--Gebelein--Rényi(HGR)最大相关函数及最小充分统计量的联系。我们刻画了一大类可学习依赖诱导表示的损失函数,包括交叉熵、合页损失及其正则化变体。特别地,我们证明这些特征可表示为损失相关函数与最大相关函数的复合,揭示了不同损失下学习表示间的本质关联。该研究还为深度分类器中观察到的神经坍缩现象提供了统计解释。最后,我们提出基于特征分离的学习设计,支持推理时进行超参数调优。
原文摘要 · Abstract (English)
We study the problem of learning feature representations from a pair of random variables, where we focus on the representations that are induced by their dependence. We provide sufficient and necessary conditions for such dependence induced representations, and illustrate their connections to Hirschfeld--Gebelein--Rényi (HGR) maximal correlation functions and minimal sufficient statistics. We characterize a large family of loss functions that can learn dependence induced representations, including cross entropy, hinge loss, and their regularized variants. In particular, we show that the features learned from this family can be expressed as the composition of a loss-dependent function and the maximal correlation function, which reveals a key connection between representations learned from different losses. Our development also gives a statistical interpretation of the neural collapse phenomenon observed in deep classifiers. Finally, we present the learning design based on the feature separation, which allows hyperparameter tuning during inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。