针对多任务线性回归,提出依赖数据分布的泛化误差新界。
Distribution-dependent Generalization Bounds for Tuning Linear Regression Across Tasks
- 基于任务数据分布特性,改进正则化系数调优的泛化分析
- 在高维下仍保持紧致界,尤其对子高斯分布效果显著
- 适合需要跨任务调参且关注高维数据建模的研究者
现代回归问题常涉及高维数据,正则化超参数的精细调优对防止过拟合、实现有效变量选择至关重要。本文研究在多个相关任务间联合调优线性回归正则化参数的新方向。针对L1、L2系数(包括岭回归、Lasso及弹性网),我们给出了验证损失的分布相关泛化误差上界。与以往适用于所有分布的统一界不同,这些界随数据分布“友好程度”提升而变紧。具体而言,在任务内样本为独立同分布且来自子高斯等经典分布类的前提下,我们的界不随特征维度d增加而恶化,且在d极大时远优于已有结果。此外,我们将结论拓展至一种广义岭回归,通过引入真实均值估计进一步收紧了界。
原文摘要 · Abstract (English)
Modern regression problems often involve high-dimensional data and a careful tuning of the regularization hyperparameters is crucial to avoid overly complex models that may overfit the training data while guaranteeing desirable properties like effective variable selection. We study the recently introduced direction of tuning regularization hyperparameters in linear regression across multiple related tasks. We obtain distribution-dependent bounds on the generalization error for the validation loss when tuning the L1 and L2 coefficients, including ridge, lasso and the elastic net. In contrast, prior work develops bounds that apply uniformly to all distributions, but such bounds necessarily degrade with feature dimension, d. While these bounds are shown to be tight for worst-case distributions, our bounds improve with the "niceness" of the data distribution. Concretely, we show that under additional assumptions that instances within each task are i.i.d. draws from broad well-studied classes of distributions including sub-Gaussians, our generalization bounds do not get worse with increasing d, and are much sharper than prior work for very large d. We also extend our results to a generalization of ridge regression, where we achieve tighter bounds that take into account an estimate of the mean of the ground truth distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。