arXiv:2602.02358stat.MLcs.LG2026-02被引 1

用分位数匹配实现跨域数据迁移,提升小样本回归预测效果

Transfer Learning Through Conditional Quantile Matching

  • 为每个源域建独立生成模型,通过分位数匹配对齐目标域分布
  • 在真实数据和模拟实验中均显著优于仅用目标域数据的模型
  • 无需假设协变量或标签分布变化,适合复杂迁移场景

我们提出一种用于回归任务的迁移学习框架,利用异构源域提升数据稀缺目标域的预测性能。方法为每个源域分别构建条件生成模型,并通过条件分位数匹配将生成结果校准至目标域分布。该分布对齐步骤在不施加协变量或标签漂移等限制性假设的前提下,纠正了源与目标域间的普遍差异。所提框架为下游目标域学习提供了原理清晰且灵活的数据增强方式。理论上,在温和条件下,基于增强数据集训练的经验风险最小化(ERM)模型比仅使用目标域数据的ERM具有更紧的过拟合界;特别地,我们推导出控制迁移偏差-方差权衡的分位数匹配估计器的新收敛速率。实践上,大量模拟实验与真实数据应用表明,该方法持续优于仅用目标域数据的学习及现有迁移学习方法。

原文摘要 · Abstract (English)

We introduce a transfer learning framework for regression that leverages heterogeneous source domains to improve predictive performance in a data-scarce target domain. Our approach learns a conditional generative model separately for each source domain and calibrates the generated responses to the target domain via conditional quantile matching. This distributional alignment step corrects general discrepancies between source and target domains without imposing restrictive assumptions such as covariate or label shift. The resulting framework provides a principled and flexible approach to high-quality data augmentation for downstream learning tasks in the target domain. From a theoretical perspective, we show that an empirical risk minimizer (ERM) trained on the augmented dataset achieves a tighter excess risk bound than the target-only ERM under mild conditions. In particular, we establish new convergence rates for the quantile matching estimator that governs the transfer bias-variance tradeoff. From a practical perspective, extensive simulations and real data applications demonstrate that the proposed method consistently improves prediction accuracy over target-only learning and competing transfer learning methods.

迁移学习回归预测分位数匹配数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。