arXiv:2412.18081stat.MLcs.LG2024-12被引 8

解决跨域回归中特征不匹配问题,实现高维数据迁移学习。

Heterogeneous transfer learning for high-dimensional regression with feature mismatch

  • 通过源域数据学习缺失特征映射,补全目标域缺失变量。
  • 理论证明误差率最优,且优于传统同质迁移学习方法。
  • 适用于多源场景,可自动排除有害源域,适合科研与医疗建模。

研究高维回归中的异构迁移学习(HTL),解决源域与目标域特征集不一致的问题。当源域某些变量在目标域不可用时,传统同质迁移学习失效。现有异构方法缺乏统计误差保证。本文提出新方法:先利用源域海量数据学习缺失特征与可观测特征的映射关系,对目标域缺失特征进行插补,再进行两阶段正则化回归迁移。考虑线性与非参数映射。在源与目标模型稀疏差异假设下,建立估计与预测误差的上界,并给出匹配的极小极大下界,证明方法达到最优率。结果揭示了模型复杂度、样本量、特征映射质量与域间差异的影响。进一步推导了忽略不可用特征的同质迁移学习的极小极大率,表明本文方法误差更小。还扩展至多源场景,设计可证明排除对抗性源域的负迁移防御机制。

原文摘要 · Abstract (English)

We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets. Such feature mismatch arises when some variables available in a data-rich source domain are unavailable in a data-poor target domain. Yet most homogeneous TL methods require the same feature space in both the source and target domains, limiting their practical applicability. Conversely, existing HTL methods lack statistical error guarantees, limiting their utility for scientific discovery. We propose an HTL method that first learns a feature map between the missing and observed features leveraging the vast source data, imputes the unavailable features in the target, and then performs a two-step TL for penalized regression. We consider both the linear and the nonparametric feature maps. We develop upper bounds on the estimation and prediction errors of HTL, assuming that the source and target parameters differ sparsely, without requiring the target model itself to be sparse. We also establish matching minimax lower bounds, showing that the proposed procedures achieve optimal rates. Our results elucidate the effects of model complexity, sample size, the quality and differences in feature maps, and differences in the models across domains. We also derive minimax rates for the misspecified homogeneous TL model that discards unavailable features and show that our HTL procedure can attain a smaller error rate than homogeneous TL. We further extend the framework to multiple source domains and develop a negative-transfer defense that provably excludes adversarial sources from transfer with high probability.

迁移学习高维回归特征匹配统计保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。