arXiv:2409.18228cs.CV2024-09中稿 · ECCV被引 1

提出基于像素距离的约束,提升自监督模型在物体中心分布上的表现

Analysis of Spatial augmentation in Self-supervised models in the purview of training and test distributions

  • 将随机裁剪拆解为重叠与补丁两部分,分析其对下游任务的影响
  • 发现裁切增强无法学习良好表征,因破坏了场景级结构信息
  • 引入像素距离相关的边距机制,有效缓解训练与测试分布差距

本文针对自监督表示学习中常用的两种空间增强方法(随机裁剪与裁切),进行实证研究。我们首次将随机裁剪分解为重叠和补丁两个独立成分,系统分析其面积对下游任务准确率的影响。同时揭示裁切增强难以学习有效表征的原因:它破坏了场景级语义一致性。基于此,我们提出一种基于距离的边距约束加入不变性损失,使模型在场景中心分布上学习更鲁棒的表征。实验表明,仅使用两个空间视图间的像素距离作为比例边距,即可显著提升性能,有助于弥合训练增强与测试分布之间的域偏移。

原文摘要 · Abstract (English)

In this paper, we present an empirical study of typical spatial augmentation techniques used in self-supervised representation learning methods (both contrastive and non-contrastive), namely random crop and cutout. Our contributions are: (a) we dissociate random cropping into two separate augmentations, overlap and patch, and provide a detailed analysis on the effect of area of overlap and patch size to the accuracy on down stream tasks. (b) We offer an insight into why cutout augmentation does not learn good representation, as reported in earlier literature. Finally, based on these analysis, (c) we propose a distance-based margin to the invariance loss for learning scene-centric representations for the downstream task on object-centric distribution, showing that as simple as a margin proportional to the pixel distance between the two spatial views in the scence-centric images can improve the learned representation. Our study furthers the understanding of the spatial augmentations, and the effect of the domain-gap between the training augmentations and the test distribution.

自监督学习空间增强表征学习域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。