arXiv:2410.21869cs.LGcs.AI2024-10ICLR被引 33

交叉熵可还原数据生成过程中的真实变量。

Cross-Entropy Is All You Need To Invert the Data Generating Process

  • 用交叉熵最小化实现数据生成过程的逆向解构。
  • 在模拟数据和ImageNet上均发现线性可解的潜在结构。
  • 为监督学习的有效性提供理论解释,适合研究者参考。

监督学习已成为现代机器学习的核心,但其有效性的完整理论仍不明确。实证现象如神经类比与线性表征假说表明,监督模型能以线性方式学习可解释的变因。近期自监督学习(特别是非线性独立成分分析)进展显示,这些方法可通过逆向数据生成过程恢复潜在结构。本文将可识别性结果扩展至参数化实例判别,并证明在标准分类任务中,交叉熵最小化可使模型学习到真实变因的线性变换表示。我们通过一系列实验验证:首先,在符合理论假设的模拟数据中成功解耦潜在因素;其次,在广泛使用的解耦基准DisLib上,简单分类任务可恢复潜在结构的线性形式;最后,在ImageNet上训练的模型能线性解码代理变因。理论与实证结合,为神经网络中超级叠加等现象提供了合理解释。本工作推动了对监督深度学习‘不合理有效性’的统一理论构建。

原文摘要 · Abstract (English)

Supervised learning has become a cornerstone of modern machine learning, yet a comprehensive theory explaining its effectiveness remains elusive. Empirical phenomena, such as neural analogy-making and the linear representation hypothesis, suggest that supervised models can learn interpretable factors of variation in a linear fashion. Recent advances in self-supervised learning, particularly nonlinear Independent Component Analysis, have shown that these methods can recover latent structures by inverting the data generating process. We extend these identifiability results to parametric instance discrimination, then show how insights transfer to the ubiquitous setting of supervised learning with cross-entropy minimization. We prove that even in standard classification tasks, models learn representations of ground-truth factors of variation up to a linear transformation. We corroborate our theoretical contribution with a series of empirical studies. First, using simulated data matching our theoretical assumptions, we demonstrate successful disentanglement of latent factors. Second, we show that on DisLib, a widely-used disentanglement benchmark, simple classification tasks recover latent structures up to linear transformations. Finally, we reveal that models trained on ImageNet encode representations that permit linear decoding of proxy factors of variation. Together, our theoretical findings and experiments offer a compelling explanation for recent observations of linear representations, such as superposition in neural networks. This work takes a significant step toward a cohesive theory that accounts for the unreasonable effectiveness of supervised deep learning.

监督学习交叉熵表征解耦理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。