arXiv:2502.09047stat.MLcs.LG2025-02被引 3

提出线性回归中协变量偏移下的最优算法,关键在于预条件变换。

Optimal Algorithms in Linear Regression under Covariate Shift: On the Importance of Precondition

  • 通过凸优化计算源与目标分布间的最优线性预条件变换。
  • 证明了在椭球约束下,最优泛化误差有紧的贝叶斯克拉美罗下界。
  • 揭示SGD及其加速版本何时能达到理论最优,适用于特定目标分布类。

现代统计学习的一个核心挑战是在源数据分布外(OOD)实现良好泛化,即使在经典的协变量偏移线性模型设定下,该问题仍未解决。本文研究高维线性回归中的基础问题:给定目标协变量矩阵,协变量偏移下的最小最大最优算法是什么?哪些目标类中常用的SGD类算法可达到最优?分析从基于贝叶斯克拉美罗不等式建立紧的泛化误差下界开始。针对问题(i),证明最优估计器仅是源分布最优估计器的某种线性变换,且该变换可通过凸规划高效求解。针对问题(ii),通过将算法累积更新与理想变换视为对学习变量的预条件,给出了SGD及其加速版本达到最优的充分条件。

原文摘要 · Abstract (English)

A common pursuit in modern statistical learning is to attain satisfactory generalization out of the source data distribution (OOD). In theory, the challenge remains unsolved even under the canonical setting of covariate shift for the linear model. This paper studies the foundational (high-dimensional) linear regression where the ground truth variables are confined to an ellipse-shape constraint and addresses two fundamental questions in this regime: (i) given the target covariate matrix, what is the min-max \emph{optimal} algorithm under covariate shift? (ii) for what kinds of target classes, the commonly-used SGD-type algorithms achieve optimality? Our analysis starts with establishing a tight lower generalization bound via a Bayesian Cramer-Rao inequality. For (i), we prove that the optimal estimator can be simply a certain linear transformation of the best estimator for the source distribution. Given the source and target matrices, we show that the transformation can be efficiently computed via a convex program. The min-max optimal analysis for SGD leverages the idea that we recognize both the accumulated updates of the applied algorithms and the ideal transformation as preconditions on the learning variables. We provide sufficient conditions when SGD with its acceleration variants attain optimality.

线性回归协变量偏移最优算法预条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。