arXiv:2511.21466cs.LGmath.OC2025-11被引 2

用平均场模型分析两层神经网络的共识优化,提升训练效率与内存表现。

Mean-Field Model for Two-Layer Neural Networks Trained with Consensus-Based Optimization

  • 将共识优化嵌入最优传输框架,构建平均场模型。
  • 粒子数无穷时,方差单调下降,均值场模型可收敛。
  • 结合Adam加速收敛,适合多任务学习场景。

本文研究两层神经网络训练中的共识优化(CBO)方法。在两个测试案例中对比CBO与Adam的表现,发现混合策略(CBO+Adam)比纯CBO收敛更快。针对多任务学习,提出一种新形式的CBO,显著降低内存开销。通过最优传输框架重构建模CBO,并在粒子数趋于无穷时,将其动力学提升至Wasserstein- over-Wasserstein空间,证明方差单调递减。数值实验验证了两种平均场模型的收敛性。

原文摘要 · Abstract (English)

We study Consensus-Based Optimization (CBO) for two-layer neural network training. We compare the performance of CBO against Adam on two test cases and demonstrate how a hybrid approach, combining CBO with Adam, provides faster convergence than CBO. Additionally, in the context of multi-task learning, we recast CBO into a formulation that offers less memory overhead. The CBO method allows for a mean-field model formulation, which we couple with the mean-field model of the neural network. To this end, we first reformulate CBO within the optimal transport framework. As the number of particles tends to infinity, we lift the corresponding dynamics to the Wasserstein-over-Wasserstein space and show that the variance decreases monotonically. We confirm numerically that both mean-field models converge.

神经网络优化算法平均场多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。