研究数据变换下的分布外泛化,给出理论保障与学习算法。
Transformation-Invariant Learning and Theoretical Guarantees for OOD Generalization
- 基于变换不变性设计学习规则,将问题转化为经验风险最小化。
- 样本复杂度上界由变换后预测器类的VC维决定,通常接近原类的VC维。
- 提出博弈视角:学习者与对手对抗最坏情况损失,适合鲁棒学习场景。
当训练与测试分布相同时,统计学习已被广泛研究。然而,在分布偏移情况下仍存在诸多未解问题。本文聚焦于训练与测试分布可通过数据变换映射关联的设置,首次对该框架进行理论分析,探讨目标变换类已知或未知的学习场景。提出学习规则与到经验风险最小化(ERM)的算法约简,并给出学习保证。建立了样本复杂度的上界,其依赖于变换后预测器类的VC维,我们证明在多数情形下该值并不显著大于原始预测器类的VC维。特别地,所推导的学习规则揭示了一种博弈论视角:学习者寻找预测器,对手寻找变换映射,分别最小化和最大化最坏情况损失。
原文摘要 · Abstract (English)
Learning with identical train and test distributions has been extensively investigated both practically and theoretically. Much remains to be understood, however, in statistical learning under distribution shifts. This paper focuses on a distribution shift setting where train and test distributions can be related by classes of (data) transformation maps. We initiate a theoretical study for this framework, investigating learning scenarios where the target class of transformations is either known or unknown. We establish learning rules and algorithmic reductions to Empirical Risk Minimization (ERM), accompanied with learning guarantees. We obtain upper bounds on the sample complexity in terms of the VC dimension of the class composing predictors with transformations, which we show in many cases is not much larger than the VC dimension of the class of predictors. We highlight that the learning rules we derive offer a game-theoretic viewpoint on distribution shift: a learner searching for predictors and an adversary searching for transformation maps to respectively minimize and maximize the worst-case loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。