提出抗概念漂移的对比预训练方法,提升动态数据下的模型鲁棒性。
Resilient Contrastive Pre-training under Non-Stationary Drift
- 基于因果干预设计新目标,缓解数据分布变化带来的偏差
- 在多个下游任务中显著降低概念漂移的影响,提升特征稳定性
- 适合处理持续变化的数据流,如在线推荐、传感器监测等场景
大规模对比预训练的成功主要依赖于规模巨大但静态的数据集。然而,随着模型规模扩大,当应用于具有概念漂移(即底层数据分布不可预测变化)的动态数据流时,该范式面临根本性挑战。本文揭示传统对比预训练方法对概念漂移高度敏感,导致学习到的特征表示出现显著偏差和不稳定性。为此,我们构建了一个结构因果模型,阐明漂移如何作为混淆因子扭曲表示。基于此因果分析,提出抗漂移对比预训练(RCP),通过因果干预设计新目标,主动消除漂移带来的偏差。RCP方法实现简单且可扩展,具备良好适应性,支持在非平稳数据上进行鲁棒、自主的预训练。跨多种下游任务的实验表明,RCP能有效缓解概念漂移的负面影响,获得更具韧性与泛化能力的表示。
原文摘要 · Abstract (English)
The remarkable success of large-scale contrastive pre-training has been largely driven by by vast yet static datasets. However, as the scaling paradigm evolves, this paradigm encounters a fundamental challenge when applied to dynamic data streams characterized by concept drift - unpredictable changes in the underlying data distribution. This paper aims to advance robust pre-training under such non-stationary environments. We begin by revealing that conventional contrastive pre-training methods are highly susceptible to concept drift, resulting in significant substantial bias and instability within the learned feature representations. To systematically analyze these effects, we develop a structural causal model that elucidates how drift acts as a confounder, distorting the learned representations. Based on these causal insights, we propose Resilient Contrastive Pre-training (RCP), a novel method that incorporates causal intervention. RCP formulates a causally-informed objective to mitigate drift-induced biases through targeted interventions. The method is designed for simple and scalable implementation and exhibits notable adaptability, promoting robust and autonomous pre-training on non-stationary data. Comprehensive experiments across various downstream tasks consistently demonstrate that RCP effectively alleviates the detrimental impact of concept drift, yielding more resilient and generalizable representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。