提出简单超参选择策略,显著降低迁移学习中调参负担
Transfer Learning in $\ell_1$ Regularized Regression: Hyperparameter Selection Strategy based on Sharp Asymptotic Analysis
- 基于复制方法的渐近分析揭示迁移规律
- 忽略一种信息传递对泛化性能影响小
- 适合高维稀疏回归中的迁移学习研究者
迁移学习旨在利用多个相关数据集的信息提升目标数据集的预测性能。在高维稀疏回归场景下,已有基于Lasso的算法如Trans-Lasso和Pretraining Lasso被提出,但其超参数(控制信息迁移程度与类型)的选择策略及对性能的影响尚未深入研究。本文通过复制方法进行高维设置下的精确渐近分析,发现一个出人意料的现象:在微调阶段忽略两种信息传递中的一种,对泛化性能影响甚微,表明超参数调优可大幅简化。该理论结果在真实数据集(IMDb)和半人工数据集(MNIST)上均得到实证支持。
原文摘要 · Abstract (English)
Transfer learning techniques aim to leverage information from multiple related datasets to enhance prediction quality against a target dataset. Such methods have been adopted in the context of high-dimensional sparse regression, and some Lasso-based algorithms have been invented: Trans-Lasso and Pretraining Lasso are such examples. These algorithms require the statistician to select hyperparameters that control the extent and type of information transfer from related datasets. However, selection strategies for these hyperparameters, as well as the impact of these choices on the algorithm's performance, have been largely unexplored. To address this, we conduct a thorough, precise study of the algorithm in a high-dimensional setting via an asymptotic analysis using the replica method. Our approach reveals a surprisingly simple behavior of the algorithm: Ignoring one of the two types of information transferred to the fine-tuning stage has little effect on generalization performance, implying that efforts for hyperparameter selection can be significantly reduced. Our theoretical findings are also empirically supported by applications on real-world and semi-artificial datasets using the IMDb and MNIST datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。