arXiv:2502.01347stat.MLcs.LG2025-02ICML被引 12

揭示高维回归中虚假相关性的形成机制与正则化的关系

Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization

  • 用协方差谱和舒尔补刻画虚假特征的简单性及其与真实特征相关性
  • 证明正则化强度λ越小,虚假相关性C越大,但测试损失L越低
  • 发现过参数化等价于正则化线性回归,解释模型偏好简单特征

在高维回归中,学习模型常依赖非预测特征与标签间的虚假相关性,影响模型鲁棒性、公平性。本文针对包含一个预测核心特征x和一个虚假特征y的数据,定量分析了线性回归中学习到的虚假相关性量C,其依赖于数据协方差和岭正则化强度λ。我们通过特征y协方差谱刻画其简单性,利用完整协方差的舒尔补描述其与x的相关性。进一步证明:使分布内测试损失L最小的λ值所处区间内,虚假相关性C随λ减小而增加,呈现权衡关系。最后通过随机特征模型研究过参数化,发现其等价于正则化线性回归。理论结果在高斯、Color-MNIST和CIFAR-10数据集上得到数值实验验证。

原文摘要 · Abstract (English)

Learning models have been shown to rely on spurious correlations between non-predictive features and the associated labels in the training data, with negative implications on robustness, bias and fairness. In this work, we provide a statistical characterization of this phenomenon for high-dimensional regression, when the data contains a predictive core feature $x$ and a spurious feature $y$. Specifically, we quantify the amount of spurious correlations $C$ learned via linear regression, in terms of the data covariance and the strength $λ$ of the ridge regularization. As a consequence, we first capture the simplicity of $y$ through the spectrum of its covariance, and its correlation with $x$ through the Schur complement of the full data covariance. Next, we prove a trade-off between $C$ and the in-distribution test loss $L$, by showing that the value of $λ$ that minimizes $L$ lies in an interval where $C$ is increasing. Finally, we investigate the effects of over-parameterization via the random features model, by showing its equivalence to regularized linear regression. Our theoretical results are supported by numerical experiments on Gaussian, Color-MNIST, and CIFAR-10 datasets.

高维回归虚假相关正则化过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。