通过融合预训练模型提升高维变量选择的微调效果,实现更优预测性能。
SMART Fine-tuning Factor Augmented Neural Lasso

- 将预训练模型作为增强特征,只学习目标任务的残差部分
- 在样本量有限时仍接近最优表现,显著优于传统微调方法
- 适合高维依赖数据与分布偏移场景下的迁移学习应用
微调是适应预训练模型到新任务的常用策略,但在高维非参数设置下进行变量选择时,其方法论和理论性质尚未完善。我们提出一种源模型增强残差调优(SMART)框架,将预训练源模型作为增强特征引入目标学习器,并仅估计目标特定的残差部分。该方法广泛适用于参数模型、稀疏模型、神经网络及黑箱机器学习模型。本文聚焦于开发因子增强神经Lasso的微调框架,即SMART-FAN-Lasso。该迁移学习框架在高维非参数回归中同时处理协变量和后验分布偏移。采用低秩因子结构应对高维相关协变量,并通过残差调优分解,将目标函数表示为源模型与其他目标特有变量的函数,从而降低目标任务的有效复杂度。我们推导出极小极大最优的超额风险界,刻画了在相对样本规模和函数复杂度条件下,微调相较于单任务学习具有统计加速的精确条件。大量数值实验在多种协变量与后验偏移场景下表明,SMART-FAN-Lasso始终优于标准基线,在极端样本受限情况下也达到近似最优性能,验证了理论速率的准确性。
原文摘要 · Abstract (English)
Fine-tuning is a widely used strategy for adapting pre-trained models to new tasks, yet its methodology and theoretical properties in high-dimensional nonparametric settings with variable selection have not yet been developed. We propose a source-model-augmented residual tuning (SMART) framework, which incorporates the pre-trained source model as an augmented feature into the target learner and estimates only the residual target-specific component. The approach is widely applicable, from parametric and sparse models to neural networks and blackbox machine learning models. We focus on the development of fine-tuning factor-augmented neural Lasso, resulting in SMART-FAN-Lasso. This transfer-learning framework for high-dimensional nonparametric regression with variable selection simultaneously handles covariate and posterior shifts. We use a low-rank factor structure to manage high-dimensional dependent covariates and a residual tuning decomposition in which the target function is expressed as a function of source model and other target-specific variables, thereby reducing the effective complexity of the target task. We derive minimax-optimal excess risk bounds, characterizing the precise conditions, in terms of relative sample sizes and function complexities, under which fine-tuning yields statistical acceleration over single-task learning. Extensive numerical experiments across diverse covariate- and posterior-shift scenarios demonstrate that SMART-FAN-Lasso consistently outperforms standard baselines and achieves near-oracle performance even under severe target sample size constraints, empirically validating the derived rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。