arXiv:2501.11265cs.LGstat.ML2025-01

用度量拓扑重新理解深度学习分类,让表现相似的模型在数学上更接近。

A Metric Topology of Deep Learning for Data Classification

  • 从概率视角定义新距离度量,使分类性能相近的模型彼此靠近。
  • 证明该度量在商空间上成立,且绝大多数网络构成紧致空间。
  • 为理解深度学习提供新数学框架,适合研究理论机制的学者。

深度学习在实际应用中表现出前所未有的成功,但其仍被视为神秘的“黑箱”,促使近期理论研究试图构建其数学基础。本文通过度量拓扑视角研究深度学习的数据分类问题。考虑到传统欧氏度量在参数空间中通常无法区分具有不同分类结果的网络,我们从概率角度提出一种有意义的距离度量,使得分类表现相似的网络在该度量下距离较近。该距离度量定义了参数向量间的等价关系:表现相同的网络属于同一等价类。有趣的是,该度量可被严格证明是商空间上的度量。在相对宽松的条件下,我们证明除了一个可忽略的子集(可能预测非唯一标签),该度量空间是紧致的,并与经典的商拓扑空间一致。本研究深化了对深度学习本质的理解,为利用丰富的度量空间理论研究深度学习开辟了新路径。

原文摘要 · Abstract (English)

Empirically, Deep Learning (DL) has demonstrated unprecedented success in practical applications. However, DL remains by and large a mysterious "black-box", spurring recent theoretical research to build its mathematical foundations. In this paper, we investigate DL for data classification through the prism of metric topology. Considering that conventional Euclidean metric over the network parameter space typically fails to discriminate DL networks according to their classification outcomes, we propose from a probabilistic point of view a meaningful distance measure, whereby DL networks yielding similar classification performances are close. The proposed distance measure defines such an equivalent relation among network parameter vectors that networks performing equally well belong to the same equivalent class. Interestingly, our proposed distance measure can provably serve as a metric on the quotient set modulo the equivalent relation. Then, under quite mild conditions it is shown that, apart from a vanishingly small subset of networks likely to predict non-unique labels, our proposed metric space is compact, and coincides with the well-known quotient topological space. Our study contributes to fundamental understanding of DL, and opens up new ways of studying DL using fruitful metric space theory.

深度学习度量拓扑理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。