arXiv:2412.03314cs.CV2024-12被引 2

通过图像重建提升自监督学习中的等变特征,增强模型泛化能力。

Equivariant Representation Learning for Augmentation-based Self-Supervised Learning via Image Reconstruction

  • 用交叉注意力融合双增强视图特征,重建其中一个视图。
  • 在3DIEBench和ImageNet上显著优于主流自监督方法。
  • 无需额外参数,适用于多种数据集与增强策略。

基于增强的自监督学习在视觉表征学习中表现优异,擅长提取不变特征,但常忽视等变特征,影响基础模型的泛化性,尤其在需等变性的下游任务中。本文提出在增强型自监督学习中引入图像重建作为辅助任务,以促进等变特征学习,且不增加额外参数。该方法通过交叉注意力机制融合两个增强视图的特征,并重建其中一个视图。该方法可适配多种数据集与基于增强对的学习方法。我们在人工数据集(3DIEBench)和自然数据集(ImageNet)上通过多线性回归任务及下游应用验证其有效性。结果一致表明,该方法显著优于标准自监督学习方法及当前先进方法,尤其在组合增强场景下表现突出。所提方法同时提升了不变与等变特征的学习能力,为计算机视觉任务提供了更鲁棒、更通用的视觉表征。

原文摘要 · Abstract (English)

Augmentation-based self-supervised learning methods have shown remarkable success in self-supervised visual representation learning, excelling in learning invariant features but often neglecting equivariant ones. This limitation reduces the generalizability of foundation models, particularly for downstream tasks requiring equivariance. We propose integrating an image reconstruction task as an auxiliary component in augmentation-based self-supervised learning algorithms to facilitate equivariant feature learning without additional parameters. Our method implements a cross-attention mechanism to blend features learned from two augmented views, subsequently reconstructing one of them. This approach is adaptable to various datasets and augmented-pair based learning methods. We evaluate its effectiveness on learning equivariant features through multiple linear regression tasks and downstream applications on both artificial (3DIEBench) and natural (ImageNet) datasets. Results consistently demonstrate significant improvements over standard augmentation-based self-supervised learning methods and state-of-the-art approaches, particularly excelling in scenarios involving combined augmentations. Our method enhances the learning of both invariant and equivariant features, leading to more robust and generalizable visual representations for computer vision tasks.

自监督学习等变特征图像重建视觉表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。