arXiv:2509.25637cs.LG2025-09

预条件通过控制特征学习的频谱偏置,提升模型泛化能力。

How Does Preconditioning Guide Feature Learning in Deep Neural Networks?

  • 用输入协方差矩阵的幂构造预条件器,控制特征学习的频谱偏好。
  • 当预条件频谱与教师模型对齐时,泛化性能显著提升。
  • 适用于关注噪声鲁棒性、分布外泛化和知识迁移的研究者。

预条件在机器学习中广泛用于加速经验风险收敛,但其对期望风险的影响仍不明确。本文研究预条件如何影响特征学习与泛化性能。我们证明,模型可获取的输入信息仅通过预条件器定义的格拉姆矩阵传递,从而引入可控的频谱偏置。具体地,在单指标教师模型中,将预条件器设为输入协方差矩阵的 $p$-次幂,发现泛化性能受指数 $p$ 及教师与输入频谱对齐程度的显著影响。从三个角度进一步分析:(i) 噪声鲁棒性,(ii) 分布外泛化,(iii) 前向知识迁移。结果表明,学习到的特征表示紧密反映预条件器引入的频谱偏置——强化被强调的成分,降低对被抑制成分的敏感性。关键发现是,当该频谱偏置与教师模型一致时,泛化性能大幅增强。

原文摘要 · Abstract (English)

Preconditioning is widely used in machine learning to accelerate convergence on the empirical risk, yet its role on the expected risk remains underexplored. In this work, we investigate how preconditioning affects feature learning and generalization performance. We first show that the input information available to the model is conveyed solely through the Gram matrix defined by the preconditioner's metric, thereby inducing a controllable spectral bias on feature learning. Concretely, instantiating the preconditioner as the $p$-th power of the input covariance matrix and within a single-index teacher model, we prove that in generalization, the exponent $p$ and the alignment between the teacher and the input spectrum are crucial factors. We further investigate how the interplay between these factors influences feature learning from three complementary perspectives: (i) Robustness to noise, (ii) Out-of-distribution generalization, and (iii) Forward knowledge transfer. Our results indicate that the learned feature representations closely mirror the spectral bias introduced by the preconditioner -- favoring components that are emphasized and exhibiting reduced sensitivity to those that are suppressed. Crucially, we demonstrate that generalization is significantly enhanced when this spectral bias is aligned with that of the teacher.

预条件特征学习泛化频谱偏置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。