arXiv:2604.02393cs.LGnlin.AO2026-04

揭示MLP训练中梯度消失与过拟合的动态机制

Plateaus, Optima, and Overfitting in Multi-Layer Perceptrons: A Saddle-Saddle-Attractor Scenario

论文配图:Plateaus, Optima, and Overfitting in Multi-Layer Perceptrons: A Saddle-Saddle-Attractor Scenario
图 1 · 摘自论文原文
  • 用极简模型解析MLP训练中的动态路径
  • 训练过程经由鞍点结构主导的平台区,最终进入过拟合
  • 小噪声数据下无法达到理论最优,必陷过拟合

梯度消失与过拟合是机器学习的核心问题,但通常在渐近条件下分析,难以揭示其动态起源。本文基于Fukumizu和Amari的启发,提出一个极简模型,描述多层感知机(MLP)的训练动态。结果显示,训练过程会经过由鞍点结构组织的平台区与近优区域,最终收敛至过拟合状态。在数据满足特定条件时,该过拟合状态坍缩为对称性下的单一吸引子。此外,对于有限噪声数据集,理论上最优解无法达到,动态过程必然陷入过拟合解。

原文摘要 · Abstract (English)

Vanishing gradients and overfitting are central problems in machine learning, yet are typically analyzed in asymptotic regimes that obscure their dynamical origins. Here we provide a dynamical description of learning in multi-layer perceptrons (MLPs) via a minimal model inspired by Fukumizu and Amari. We show that training dynamics traverse plateau and near-optimal regions, both organized by saddle structures, before converging to an overfitting regime. Under suitable conditions on the data, this regime collapses to a single attractor modulo symmetry. Furthermore, for finite noisy datasets, convergence to the theoretical optimum is impossible, and the dynamics necessarily settle into an overfitting solution.

深度学习梯度消失过拟合动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。