揭示机器学习中约束学习对经典数据处理不等式的失效,并提出新不等式。
Comparing Corrupted Constrained Learning Problems
- 在约束模型下重构数据处理不等式,引入受限贝叶斯风险
- 证明新不等式等价于超预测集的集合包含关系
- 为信息瓶颈与特征学习提供理论修正,适合理论研究者
统计学中的经典数据处理不等式表明,通过随机变换得到的统计实验的贝叶斯风险不会低于原实验。该结论在机器学习中支撑信息瓶颈原理与特征学习等应用。然而,机器学习问题本质是受限学习:模型类不包含所有可测函数。本文给出一个简单反例,证明经典不等式在此设定下不再成立。为此,我们提出广义数据处理不等式,要求联合分布的受限贝叶斯风险(相对于损失函数与受限假设类)下界住随机修改后分布的受限贝叶斯风险,且不依赖于分布选择。我们证明该不等式等价于由损失函数与模型类诱导的超预测集的集合包含条件,并推导出满足该条件的充分条件。
原文摘要 · Abstract (English)
A key result in statistics is the data processing inequality, originally proved by Blackwell (1951) and later refined by DeGroot (1962) in terms of statistical uncertainty. It states that the Bayes risk of a statistical experiment obtained by stochastically modifying another experiment cannot be lower than the Bayes risk of the original experiment, regardless of the loss function or prior chosen. In machine learning, this result underlies applications such as the information bottleneck principle and some feature learning techniques. However, machine learning problems are constrained learning problems: the model class used does not include all measurable functions. We present a simple counterexample showing that the classical data processing inequality fails to hold in such a setting. Hence, we formulate a generalized data processing inequality, requiring the constrained Bayes risk of a joint distribution (with respect to a loss function and a constrained hypothesis class) to lower bound the constrained Bayes risk on the stochastically modified distribution, regardless of the choice of distribution. We show this inequality to be equivalent to a set containment condition on a specific function set induced by the loss and model class, called the superprediction set. Finally, we derive sufficient conditions for this containment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。