arXiv:2606.22775stat.MEcs.LG2026-06

利用目标分布信息提升分布偏移下的回归精度,提出高效可扩展的改进方法。

Target-Aware Linear Regression Under Distribution Shift

论文配图:Target-Aware Linear Regression Under Distribution Shift
图 1 · 摘自论文原文
  • 结合目标变量与协变量的边缘分布,设计混合损失估计器
  • 在高信噪比下,两阶段方法几乎逼近最优性能且计算成本低
  • 理论推导并验证三类估计器在不同场景下的表现差异

训练与部署之间的分布偏移是现代AI系统面临的普遍挑战。在许多情况下,协变量和响应的目标边缘分布可通过群体观测、边界条件、模拟器配置属性或对齐时的分布约束获知,这些信息可为回归估计提供有价值的辅助。本文在多变量线性回归框架下研究该问题,假设源域与目标域的条件均值 $E[Yig|X]$ 稳定,并识别出联合使用两个目标边缘分布的混合损失估计器作为目标感知基准。然而,其直接计算需解耦合非线性优化,大规模下代价高昂。本文主要贡献是提出并评估两种计算上可行的替代方案:约束矩匹配估计器和在普通最小二乘基础上增加校准步骤的两阶段估计器。针对三类估计器,我们推导并比较了闭式渐近均方误差,揭示了可使近似方案达到或接近基准的条件,以及不成立的情形。通过三种受控偏移场景的蒙特卡洛实验,验证了理论结果,分析了三者间的准确率-运行时间权衡,并为估计器选择提供指导。特别地,两阶段估计器在高信噪比下几乎匹配混合基准,且几乎无额外开销,为非线性设置中的经验观察提供了理论支撑。

原文摘要 · Abstract (English)

Distribution shift between training and deployment is a pervasive challenge for modern AI systems. In many cases, the target marginals of covariates and response are known or specified through population-level observations, boundary conditions, properties of simulator configurations, or alignment-time distributional constraints. Such knowledge may provide valuable side information for regression estimation. We study this problem in the multivariate linear regression setting with a stable conditional mean $E[Y\mid X]$ across source and target, and identify the hybrid-loss estimator, which jointly incorporates both target marginals, as a benchmark target-aware estimator. Its direct computation, however, requires solving a coupled nonlinear optimization that is expensive at scale. Our main contribution is to develop and evaluate two computationally tractable alternatives: a constrained moment-matching estimator and a two-stage estimator that augments ordinary least squares with a calibration step. For all three estimators, we derive and compare closed-form asymptotic mean squared errors, yielding conditions under which the tractable alternatives match or closely approximate the hybrid benchmark, and regimes in which they do not. Monte Carlo experiments across three controlled shift regimes validate the theoretical results, investigate the accuracy-runtime tradeoffs among the three estimators, and translate into guidance on estimator choice. In particular, the two-stage estimator nearly matches the hybrid benchmark in the high signal-to-noise regime at essentially no additional cost, providing theoretical grounding for empirical observations in nonlinear settings.

分布偏移线性回归目标感知估计器设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。