通过随机排列提升非光滑非凸优化收敛性,理论与实证双验证。
Permutation Randomization on Nonsmooth Nonconvex Optimization: A Theoretical and Experimental Study
- 用随机排列打乱梯度更新顺序,打破收敛瓶颈。
- 在足够迭代下可逼近全局最优,且不降低原优化器收敛速度。
- 适合深度学习训练和噪声目标优化,具广泛适用性。
尽管基于梯度的随机化优化器在复杂优化任务中表现优异,但其理论基础仍不充分。本文聚焦于维度无关的非光滑非凸优化中随机化的作用,以随机排列为例进行理论与实验研究。理论上,分析表明随机排列能破坏梯度优化器的收缩行为,使算法在足够多迭代下持续收敛至全局最优;同时证明其可保持原优化器的收敛速率。实验上,对比三种基准方法,在堆叠结构神经网络训练和噪声目标函数优化任务中均验证了该方法的有效性。结果既支持理论发现,也凸显其实际优势。本工作为随机化技术的扩展分析提供了理论与实证基础。
原文摘要 · Abstract (English)
While gradient-based optimizers that incorporate randomization often showcase superior performance on complex optimization, the theoretical foundations underlying this superiority remain insufficiently understood. A particularly pressing question has emerged: What is the role of randomization in dimension-free nonsmooth nonconvex optimization? To address this gap, we investigate the theoretical and empirical impact of permutation randomization within gradient-based optimization frameworks, using it as a representative case to explore broader implications. From a theoretical perspective, our analyses reveal that permutation randomization disrupts the shrinkage behavior of gradient-based optimizers, facilitating continuous convergence toward the global optimum given a sufficiently large number of iterations. Additionally, we prove that permutation randomization can preserve the convergence rate of the underlying optimizer. On the empirical side, we conduct extensive numerical experiments comparing permutation-randomized optimizer against three baseline methods. These experiments span tasks such as training deep neural networks with stacked architectures and optimizing noisy objective functions. The results not only corroborate our theoretical insights but also highlight the practical benefits of permutation randomization. In summary, this work delivers both rigorous theoretical justification and compelling empirical evidence for the effectiveness of permutation randomization. Our findings and evidence lay a foundation for extending analytics to encompass a wide array of randomization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。