arXiv:2509.02407cs.LGphysics.data-an2025-09被引 4

用费舍尔信息流指导神经网络训练,避免过拟合

Fisher information flow in artificial neural networks

  • 通过追踪费舍尔信息在神经网络中的传播路径
  • 发现最优性能对应信息传输最大,过训则导致信息丢失
  • 无需验证集即可自动停止训练,适合物理测量场景

从测量数据中估计连续参数在物理学诸多领域具有核心意义。理解与改进此类估计过程的关键工具是费舍尔信息,它量化了未知参数信息在物理系统中的传播,并决定精度的理论极限。随着人工神经网络(ANN)逐渐成为许多测量系统的核心组件,理解其内部如何处理和传递与参数相关的信息变得至关重要。本文提出一种方法,用于监控执行参数估计任务的神经网络中费舍尔信息的流动,从输入层追踪至输出层。我们证明,最优估计性能对应于费舍尔信息的最大传输;超过该点继续训练将因过拟合导致信息损失。这提供了一种无需依赖独立验证数据集的模型无关训练终止准则。为展示该方法的实际意义,我们将其应用于一个基于成像实验数据训练的网络,在真实的物理设置中验证了其有效性。

原文摘要 · Abstract (English)

The estimation of continuous parameters from measured data plays a central role in many fields of physics. A key tool in understanding and improving such estimation processes is the concept of Fisher information, which quantifies how information about unknown parameters propagates through a physical system and determines the ultimate limits of precision. With Artificial Neural Networks (ANNs) gradually becoming an integral part of many measurement systems, it is essential to understand how they process and transmit parameter-relevant information internally. Here, we present a method to monitor the flow of Fisher information through an ANN performing a parameter estimation task, tracking it from the input to the output layer. We show that optimal estimation performance corresponds to the maximal transmission of Fisher information, and that training beyond this point results in information loss due to overfitting. This provides a model-free stopping criterion for network training-eliminating the need for a separate validation dataset. To demonstrate the practical relevance of our approach, we apply it to a network trained on data from an imaging experiment, highlighting its effectiveness in a realistic physical setting.

神经网络信息流参数估计过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。