arXiv:2410.10773cs.LGcs.CV2024-10被引 1

通过空间条件化提升IJEPA的表征能力与鲁棒性

Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning

  • 在IJEPA中引入位置条件编码,使编码器自适应调节预测特征
  • 在多个图像分类数据集上实现性能提升,预训练样本效率更高
  • 对上下文窗口大小变化更鲁棒,适合资源受限场景

基于图像的联合嵌入预测架构(IJEPA)为利用掩码图像建模框架进行表征学习提供了一种有吸引力的替代方案。IJEPA通过在潜在空间而非输入空间进行预测,促使表征捕捉有用语义信息。然而,IJEPA依赖精心设计的上下文窗口和目标窗口以避免表征坍缩。IJEPA中的编码器模块无法根据掩码预测任务的可行性自适应调节预测或目标特征类型,因缺乏对上下文与目标的充分信息。基于自然图像中信息具有强空间偏置的直觉——局部区域彼此高度可预测,而远距离区域则不然——我们分别在上下文和目标编码器中引入上下文与目标窗口的位置信息作为条件。实验表明,这种“条件化”编码器在多个图像分类基准数据集上均取得性能提升,同时增强了对上下文窗口大小变化的鲁棒性,并在预训练阶段提升了样本效率。

原文摘要 · Abstract (English)

Image-based Joint-Embedding Predictive Architecture (IJEPA) offers an attractive alternative to Masked Autoencoder (MAE) for representation learning using the Masked Image Modeling framework. IJEPA drives representations to capture useful semantic information by predicting in latent rather than input space. However, IJEPA relies on carefully designed context and target windows to avoid representational collapse. The encoder modules in IJEPA cannot adaptively modulate the type of predicted and/or target features based on the feasibility of the masked prediction task as they are not given sufficient information of both context and targets. Based on the intuition that in natural images, information has a strong spatial bias with spatially local regions being highly predictive of one another compared to distant ones. We condition the target encoder and context encoder modules in IJEPA with positions of context and target windows respectively. Our "conditional" encoders show performance gains on several image classification benchmark datasets, improved robustness to context window size and sample-efficiency during pretraining.

表征学习视觉模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。