arXiv:2504.18212stat.MLcs.LG2025-04被引 1

为迁移学习中的高维回归提供可验证的特征显著性检验方法。

Post-Transfer Learning Statistical Inference in High-Dimensional Regression

  • 基于分治策略构建后迁移学习统计推断框架
  • 在α=0.05下严格控制假阳性率,支持可靠特征选择
  • 适用于小样本高维场景,适合需要可解释性的研究者

迁移学习(TL)在高维回归(HDR)中具有重要意义,尤其在目标任务样本量有限时。然而,当前尚无方法量化特征与响应变量之间关系的统计显著性。本文提出一种新的统计推断框架PTL-SI(Post-TL Statistical Inference),用于评估TL-HDR中特征选择的可靠性。其核心贡献是为TL-HDR中选出的特征提供有效的$ p $-值,从而在设定显著性水平α(如0.05)下严格控制假阳性率(FPR)。此外,通过引入分治策略提升统计功效。我们在合成数据和真实高维数据集上进行了大量实验,验证了PTL-SI的有效性与理论性质,证实其在测试迁移学习中特征选择可靠性方面的实用性。

原文摘要 · Abstract (English)

Transfer learning (TL) for high-dimensional regression (HDR) is an important problem in machine learning, particularly when dealing with limited sample size in the target task. However, there currently lacks a method to quantify the statistical significance of the relationship between features and the response in TL-HDR settings. In this paper, we introduce a novel statistical inference framework for assessing the reliability of feature selection in TL-HDR, called PTL-SI (Post-TL Statistical Inference). The core contribution of PTL-SI is its ability to provide valid $p$-values to features selected in TL-HDR, thereby rigorously controlling the false positive rate (FPR) at desired significance level $α$ (e.g., 0.05). Furthermore, we enhance statistical power by incorporating a strategic divide-and-conquer approach into our framework. We demonstrate the validity and effectiveness of the proposed PTL-SI through extensive experiments on both synthetic and real-world high-dimensional datasets, confirming its theoretical properties and utility in testing the reliability of feature selection in TL scenarios.

迁移学习高维回归统计推断特征选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。