arXiv:2509.06308stat.MLcs.LG2025-09

在高维加性回归中,提出最优迁移学习方法,适应重尾误差分布。

Minimax optimal transfer learning for high-dimensional additive regression

  • 基于局部线性平滑的平滑反向拟合估计器,处理高维加性回归。
  • 在亚指数误差下达到极小极大最优率,突破传统证明方法。
  • 两阶段迁移学习框架,适用于辅助数据与目标数据相近的情形。

本文研究在迁移学习框架下的高维加性回归问题,其中除目标总体样本外,还观测到来自不同但可能相关的回归模型的辅助样本。首先提出一种仅基于目标数据的估计方法,采用带局部线性平滑的平滑反向拟合估计器。不同于以往工作,该方法在亚魏布尔(α)噪声条件下建立了通用误差界,可容纳重尾误差分布。在亚指数情形(α=1)下,证明该估计器在正则条件下达到极小极大下界,需显著偏离现有证明策略。随后,在迁移学习框架内提出一种新颖的两阶段估计方法,并在总体与经验层面提供理论保证。在一般尾部条件下推导各阶段的误差界,进一步证明当辅助分布与目标分布足够接近时,可实现极小极大最优率。所有理论结果均通过模拟研究和真实数据分析验证。

原文摘要 · Abstract (English)

This paper studies high-dimensional additive regression under the transfer learning framework, where one observes samples from a target population together with auxiliary samples from different but potentially related regression models. We first introduce a target-only estimation procedure based on the smooth backfitting estimator with local linear smoothing. In contrast to previous work, we establish general error bounds under sub-Weibull($α$) noise, thereby accommodating heavy-tailed error distributions. In the sub-exponential case ($α=1$), we show that the estimator attains the minimax lower bound under regularity conditions, which requires a substantial departure from existing proof strategies. We then develop a novel two-stage estimation method within a transfer learning framework, and provide theoretical guarantees at both the population and empirical levels. Error bounds are derived for each stage under general tail conditions, and we further demonstrate that the minimax optimal rate is achieved when the auxiliary and target distributions are sufficiently close. All theoretical results are supported by simulation studies and real data analysis.

高维回归迁移学习极小极大最优加性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。