arXiv:2501.07400cs.LGcs.AI2025-01

揭示ReLU神经网络梯度流的动态剪裁机制,解释数据复杂性如何随训练指数下降。

Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning

  • 基于欧氏损失与坐标系自适应假设,推导出权重累积的显式梯度流方程。
  • 数据聚类在输入层以指数速率被动态剪裁,剪裁速度随已处理数据量增加而加快。
  • 为监督学习的可解释性提供新视角,适合研究模型内部机制的学者。

本文基于输入层的欧氏损失梯度下降,结合ReLU激活函数与权重对激活所定义坐标系的精确适配假设,推导出深度学习中累积偏差和权重的显式方程。研究表明,梯度下降对应于输入层的动态过程:数据聚类逐步降低复杂性(“剪裁”),且剪裁速率呈指数增长,随已剪裁数据点数量增加而加快。文中详细讨论了梯度流方程的多种解类型。本工作的主要动机在于揭示监督学习中的可解释性问题。

原文摘要 · Abstract (English)

We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations. We show that gradient descent corresponds to a dynamical process in the input layer, whereby clusters of data are progressively reduced in complexity ("truncated") at an exponential rate that increases with the number of data points that have already been truncated. We provide a detailed discussion of several types of solutions to the gradient flow equations. A main motivation for this work is to shed light on the interpretability question in supervised learning.

深度学习梯度流可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。