arXiv:2502.15540stat.MLcs.IT2025-02ICLR被引 4

用数据依赖的高斯混合先验提升表示学习泛化能力

Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture Priors

  • 基于相对熵构建泛化误差上界,结合训练与测试数据的描述长度
  • 新方法在多个数据集上超越VIB和CDVIB,提升显著
  • 自动学习数据相关先验并生成注意力机制,适合表示学习研究者

我们建立了表示学习类算法的期望和尾部泛化误差界,其形式为从训练与“测试”数据集中提取的表示分布与一个数据依赖的对称先验(即潜变量的最小描述长度)之间的相对熵。该界能反映编码器的“结构”与“简洁性”,显著优于现有少数针对该模型的边界。随后,我们利用期望界设计出合适的数据依赖正则项;深入探讨了先验选择问题,提出一种系统性方法,同时学习数据依赖的高斯混合先验并将其作为正则项使用。有趣的是,该过程自然涌现出加权注意力机制。实验表明,本方法在多个数据集上优于当前流行的变分信息瓶颈(VIB)及近期的类别依赖变分信息瓶颈(CDVIB)。

原文摘要 · Abstract (English)

We establish in-expectation and tail bounds on the generalization error of representation learning type algorithms. The bounds are in terms of the relative entropy between the distribution of the representations extracted from the training and "test'' datasets and a data-dependent symmetric prior, i.e., the Minimum Description Length (MDL) of the latent variables for the training and test datasets. Our bounds are shown to reflect the "structure" and "simplicity'' of the encoder and significantly improve upon the few existing ones for the studied model. We then use our in-expectation bound to devise a suitable data-dependent regularizer; and we investigate thoroughly the important question of the selection of the prior. We propose a systematic approach to simultaneously learning a data-dependent Gaussian mixture prior and using it as a regularizer. Interestingly, we show that a weighted attention mechanism emerges naturally in this procedure. Our experiments show that our approach outperforms the now popular Variational Information Bottleneck (VIB) method as well as the recent Category-Dependent VIB (CDVIB).

表示学习泛化保证高斯混合正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。