arXiv:2412.20724stat.MLcs.LG2024-12

用新型软菱形正则化提升深度模型精度与稀疏性,优于传统方法。

Soft Diamond Regularizers for Deep Learning

  • 基于厚尾α稳定分布设计软菱形正则化,增强权重搜索能力。
  • 在CIFAR、Caltech和IWSLT数据集上提升准确率并增加权重稀疏度。
  • 适合追求高稀疏性与强泛化的大型机器学习任务使用。

本文提出一类基于厚尾对称α稳定(SαS)概率钟形曲线的软菱形突触正则化方法。这些参数化权重先验在图像和语言翻译测试集上提升了深度学习性能,并增强了训练权重的稀疏性。相比稀疏套索回归中的硬菱形拉普拉斯正则化,该方法表现更优。SαS先验具有幂律尾部,比岭回归所依赖的高斯分布指数尾部更厚,且α参数越小尾部越厚,能建模更强烈的脉冲行为,支持高维突触空间中远距离探索。其约束集几何形状呈菱形,随α尾部厚度与分布差异变化。由于SαS分布通常无闭式解,直接训练计算成本高,本文通过预计算查找表解决此瓶颈。在CIFAR-10、CIFAR-100、Caltech-256图像数据集及IWSLT-2016德英翻译数据集上的实验表明,该正则化显著提升分类器准确率与稀疏性,优于岭和套索正则化。研究推荐α=0.5的亚柯西软菱形正则化作为大规模机器学习中具有竞争力的稀疏正则化工具。

原文摘要 · Abstract (English)

This chapter presents the new family of soft diamond synaptic regularizers based on thick-tailed symmetric alpha stable $SαS$ probability bell curves. These new parametrized weight priors improved deep-learning performance on image and language-translation test sets and increased the sparsity of the trained weights. They outperformed the state-of-the-art hard-diamond Laplacian regularizer of sparse lasso regression and classification. The $SαS$ synaptic weight priors have power-law bell-curve tails that are thicker than the thin exponential tails of Gaussian bell curves that underly ridge regularizers. Their tails get thicker as the $α$ parameter decreases. These thicker tails model more impulsive behavior and allow for occasional distant search in synaptic weight spaces of extremely high dimension. The geometry of their constraint sets has a diamond shape. The shape varies from a circle to a star or diamond that depends on the $α$ tail thickness and dispersion of the $SαS$ weight prior. These $SαS$ bell curves lack a closed form in general and this makes direct training computationally intensive. We removed this computational bottleneck by using a precomputed look-up table. We tested the soft diamond regularizers with deep neural classifiers on both image test sets and German-to-English language translation. The image simulations used the three datasets CIFAR-10, CIFAR-100, and Caltech-256. The regularizers improved the accuracy and sparsity of the classifiers. We also tested with deep neural machine-translation models on the IWSLT-2016 Evaluation dataset for German-to-English text translation. They also outperformed ridge regularizers and lasso regularizers. These findings recommend the sub-Cauchy $α= 0.5$ soft diamond regularizer as a competitive and sparse regularizer for large-scale machine learning.

正则化深度学习稀疏性SαS分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。