用激活信号提前判断神经网络训练好坏,无需标签。
OUI as a Structural Observable: Towards an Activation-Centric View of Neural Network Training
- 以激活行为为观测点,提出新指标OUI判断训练质量。
- OUI能提前预判权重衰减和学习率配置是否有效。
- 适合关注训练动态、模型调试的研究者使用。
激活函数赋予深度网络表达能力,但训练过程仍主要通过损失、准确率等外部指标评估,内部结构演化鲜被观察。本文主张将过拟合-欠拟合指示器(OUI)视为网络内部结构的首个实用可观测量。在多项实验中,OUI作为早期、无标签、基于激活的信号,可提前揭示网络是否进入不良或有潜力的训练状态:在监督学习中预判权重衰减效果,在强化学习中早期区分PPO的学习除数配置,在在线控制中驱动逐层权重衰减自适应。结合近期发现——激活模式趋于稳定早于参数变化——这些结果提示了一条新研究方向:以激活为中心的训练动力学理论。OUI正成为这一理论的实证基础。
原文摘要 · Abstract (English)
Activation functions are what make deep networks expressive: without them, the model collapses to a linear map. Yet we still evaluate training mostly from the outside, through loss, accuracy, return, or final calibration, while the internal structural evolution of the network remains largely unobserved. In this paper, we argue that the Overfitting--Underfitting Indicator (OUI) should be understood as a first practical observable of that internal structure. Across our recent results, OUI consistently appears as an early, label-free, activation-based signal that reveals whether a network is entering a poor or promising training regime before convergence. In supervised learning, it anticipates weight decay regimes; in reinforcement learning, it discriminates learning-rate regimes early in PPO actor--critic; and in online control, it can drive layer-wise weight decay adaptation. Read together with recent evidence that activation patterns tend to stabilize earlier than parameters, these results suggest a broader research direction: an activation-centric theory of training dynamics. OUI is becoming an empirical foothold toward this theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。