arXiv:2502.04121cs.LGcond-mat.dis-nn2025-02被引 8

用首次通过理论优化模型训练扰动,提升效率与泛化能力。

First-Passage Approach to Optimizing Perturbations for Improved Training of Machine Learning Models

  • 将训练扰动建模为首次通过过程,预测不同频率下的响应。
  • 在CIFAR-10上找到有效扰动与频率,显著改善训练性能。
  • 方法可迁移至不同数据集、模型和任务,通用性强。

机器学习模型在物理科学应用中不可或缺,但其训练耗时远超推理时间。已有如收缩-扰动、热重启和随机重置等扰动策略可加速训练或提升泛化性,但设计多依赖直觉与试错。本文将训练过程视为首次通过过程,研究其对扰动的响应。若无扰动训练达到准稳态,单频扰动响应即可预测宽频范围行为。以ResNet-18在CIFAR-10上的分类任务为例,成功识别出有效扰动与频率。该方法进一步验证了在其他数据集、架构、优化器甚至回归任务中的可迁移性。本工作为优化训练扰动提供了理论框架。

原文摘要 · Abstract (English)

Machine learning models have become indispensable tools in applications across the physical sciences. Their training is often time-consuming, vastly exceeding the inference timescales. Several protocols have been developed to perturb the learning process and improve the training, such as shrink and perturb, warm restarts, and stochastic resetting. For classifiers, these perturbations have been shown to result in enhanced speedups or improved generalization. However, the design of such perturbations is usually done ad hoc by intuition and trial and error. To rationally optimize training protocols, we frame them as first-passage processes and consider their response to perturbations. We show that if the unperturbed learning process reaches a quasi-steady state, the response at a single perturbation frequency can predict the behavior at a wide range of frequencies. We employ this approach to a CIFAR-10 classifier using the ResNet-18 model and identify a useful perturbation and frequency among several possibilities. We demonstrate the transferability of the approach to other datasets, architectures, optimizers and even tasks (regression instead of classification). Our work allows optimization of perturbations for improving the training of machine learning models using a first-passage approach.

训练优化首次通过机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。