提出新采样策略提升自监督学习的分布外泛化能力
On the Out-of-Distribution Generalization of Self-Supervised Learning
- 基于因果模型设计后干预分布,使伪相关变量与标签独立
- 理论证明满足该分布时模型可达到最优最坏情况下的分布外性能
- 通过学习隐变量模型实现采样约束,适用于多种下游任务
本文聚焦自监督学习(SSL)的分布外(OOD)泛化问题。通过分析训练阶段的批处理构建,我们首次为SSL具备OOD泛化能力提供了合理解释。随后,从数据生成与因果推断视角出发,分析并得出结论:SSL在训练过程中会学习到伪相关性,从而降低其OOD泛化性能。为解决此问题,我们提出基于结构因果模型的后干预分布(PID),其保证伪相关变量与标签变量相互独立。此外,我们证明若每个训练批满足PID,所得SSL模型可实现最优最坏情况下的OOD表现。这一发现启发我们设计一种批采样策略,通过学习隐变量模型来强制施加PID约束。理论分析验证了隐变量模型的可识别性,并通过多个下游OOD任务的实验,充分证明了所提采样策略的有效性。
原文摘要 · Abstract (English)
In this paper, we focus on the out-of-distribution (OOD) generalization of self-supervised learning (SSL). By analyzing the mini-batch construction during the SSL training phase, we first give one plausible explanation for SSL having OOD generalization. Then, from the perspective of data generation and causal inference, we analyze and conclude that SSL learns spurious correlations during the training process, which leads to a reduction in OOD generalization. To address this issue, we propose a post-intervention distribution (PID) grounded in the Structural Causal Model. PID offers a scenario where the spurious variable and label variable is mutually independent. Besides, we demonstrate that if each mini-batch during SSL training satisfies PID, the resulting SSL model can achieve optimal worst-case OOD performance. This motivates us to develop a batch sampling strategy that enforces PID constraints through the learning of a latent variable model. Through theoretical analysis, we demonstrate the identifiability of the latent variable model and validate the effectiveness of the proposed sampling strategy. Experiments conducted on various downstream OOD tasks demonstrate the effectiveness of the proposed sampling strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。