多任务学习中,低内在维度可带来更强泛化能力。
From Low Intrinsic Dimensionality to Non-Vacuous Generalization Bounds in Deep Multi-Task Learning
- 用随机展开将多任务网络直接参数化在低维空间
- 仅需更少自由参数即可达到高精度,优于单任务学习
- 首次获得深度多任务网络的非平凡泛化界,适合理论研究者
深度学习在过参数化条件下仍能良好泛化,原因之一是其内在维度远低于环境维度(即参数量)。本文在深度多任务学习场景中验证了这一现象,提出通过随机展开技术将多任务网络直接参数化于低维空间。实验表明,高精度的多任务解可在远低于单任务学习所需内在维度下实现。进一步结合权重压缩与PAC-Bayesian分析,首次为深度多任务网络建立了非平凡的泛化界。
原文摘要 · Abstract (English)
Deep learning methods are known to generalize well from training to future data, even in an overparametrized regime, where they could easily overfit. One explanation for this phenomenon is that even when their *ambient dimensionality*, (i.e. the number of parameters) is large, the models' *intrinsic dimensionality* is small; specifically, their learning takes place in a small subspace of all possible weight configurations. In this work, we confirm this phenomenon in the setting of *deep multi-task learning*. We introduce a method to parametrize multi-task network directly in the low-dimensional space, facilitated by the use of *random expansions* techniques. We then show that high-accuracy multi-task solutions can be found with much smaller intrinsic dimensionality (fewer free parameters) than what single-task learning requires. Subsequently, we show that the low-dimensional representations in combination with *weight compression* and *PAC-Bayesian* reasoning lead to the *first non-vacuous generalization bounds* for deep multi-task networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。