arXiv:2507.17748cs.LGcs.AI2025-07ICCV被引 2

大学习率同时提升模型鲁棒性与压缩性,缓解伪相关问题。

Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility

  • 采用大学习率训练,自然促进特征不变性与稀疏激活。
  • 在多个数据集上验证,大学习率显著降低伪相关依赖,提升模型可压缩性。
  • 适合关注模型效率与鲁棒性的研究者,尤其对部署优化有需求。

鲁棒性与资源效率是现代机器学习模型的两个重要目标,但二者协同实现仍具挑战。本文发现大学习率能同时促进模型对伪相关关系的鲁棒性与网络压缩性。实验表明,大学习率可生成有利的表示特性,如不变特征利用、类别分离与激活稀疏性。相比其他超参数与正则化方法,大学习率在多种任务中更一致地满足这些性质。我们在多个存在伪相关的数据集、模型与优化器上验证了其有效性,并发现大学习率在标准分类任务中的成功,可能源于其对训练集中隐藏/罕见伪相关关系的缓解作用。机制分析揭示:在大学习率下,对偏差冲突样本的高置信误判起关键作用。

原文摘要 · Abstract (English)

Robustness and resource-efficiency are two highly desirable properties for modern machine learning models. However, achieving them jointly remains a challenge. In this paper, we identify high learning rates as a facilitator for simultaneously achieving robustness to spurious correlations and network compressibility. We demonstrate that large learning rates also produce desirable representation properties such as invariant feature utilization, class separation, and activation sparsity. Our findings indicate that large learning rates compare favorably to other hyperparameters and regularization methods, in consistently satisfying these properties in tandem. In addition to demonstrating the positive effect of large learning rates across diverse spurious correlation datasets, models, and optimizers, we also present strong evidence that the previously documented success of large learning rates in standard classification tasks is related to addressing hidden/rare spurious correlations in the training dataset. Our investigation of the mechanisms underlying this phenomenon reveals the importance of confident mispredictions of bias-conflicting samples under large learning rates.

学习率鲁棒性压缩性伪相关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。