为线性自编码器提供理论泛化边界,解释其推荐系统中的优秀表现
PAC-Bayes Bounds for Multivariate Linear Regression and Linear Autoencoders
- 基于PAC-Bayes框架推导多输出线性回归的泛化界
- 证明线性自编码器在松弛误差下可视为受限线性模型,适用该界
- 提出高效优化方法,使大模型在真实数据上可实用评估
线性自编码器(LAEs)在顶尖推荐系统中表现出色,但其成功主要依赖经验,缺乏理论支撑。本文研究多变量线性回归与LAEs的泛化能力——统计学习中衡量模型性能的理论指标。首先,我们为多变量线性回归提出了PAC-Bayes边界,扩展了Shalaeva等人针对单输出线性回归的早期结果,并建立了其收敛的充分条件。随后,我们表明,当在松弛均方误差下评估时,LAEs可被解释为在有界数据上的受限多变量线性回归模型,从而可应用本研究的边界。此外,我们发展了改进计算效率的理论方法,使得该边界能在大规模模型和真实数据集上实际计算。实验结果表明,该边界紧致且与实际排名指标如Recall@K和NDCG@K具有良好相关性。
原文摘要 · Abstract (English)
Linear Autoencoders (LAEs) have shown strong performance in state-of-the-art recommender systems. However, this success remains largely empirical, with limited theoretical understanding. In this paper, we investigate the generalizability -- a theoretical measure of model performance in statistical learning -- of multivariate linear regression and LAEs. We first propose a PAC-Bayes bound for multivariate linear regression, extending the earlier bound for single-output linear regression by Shalaeva et al., and establish sufficient conditions for its convergence. We then show that LAEs, when evaluated under a relaxed mean squared error, can be interpreted as constrained multivariate linear regression models on bounded data, to which our bound adapts. Furthermore, we develop theoretical methods to improve the computational efficiency of optimizing the LAE bound, enabling its practical evaluation on large models and real-world datasets. Experimental results demonstrate that our bound is tight and correlates well with practical ranking metrics such as Recall@K and NDCG@K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。