提出适用于实值无界损失的新压缩泛化界,突破传统限制。
Sample Compression Unleashed: New Generalization Bounds for Real Valued Losses
- 用P2L元算法将任意模型转为可压缩形式
- 首次实现对实值无界损失的紧泛化界
- 适合研究泛化理论与深度学习的从业者
样本压缩理论为仅依赖训练数据子集和简短消息串(通常为二进制序列)即可定义的预测器提供泛化保证。以往工作仅针对零一损失给出泛化界,该设定在应用于深度学习时存在明显局限。本文提出一个新框架,推导出适用于实值无界损失的样本压缩泛化界。通过使用Pick-To-Learn(P2L)元算法,可将任意机器学习预测器的训练过程转换为生成样本压缩预测器的方法。我们在随机森林及多种神经网络上进行了实验,验证了该界在不同场景下的紧致性与通用性。
原文摘要 · Abstract (English)
The sample compression theory provides generalization guarantees for predictors that can be fully defined using a subset of the training dataset and a (short) message string, generally defined as a binary sequence. Previous works provided generalization bounds for the zero-one loss, which is restrictive notably when applied to deep learning approaches. In this paper, we present a general framework for deriving new sample compression bounds that hold for real-valued unbounded losses. Using the Pick-To-Learn (P2L) meta-algorithm, which transforms the training method of any machine-learning predictor to yield sample-compressed predictors, we empirically demonstrate the tightness of the bounds and their versatility by evaluating them on random forests and multiple types of neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。