arXiv:2506.18104cs.CVcs.LG2025-06

改进VICReg模型,提升对未知数据的泛化能力与全局语义捕捉。

Enhancing VICReg: Random-Walk Pairing for Improved Generalization and Better Global Semantics Capturing

  • 引入随机游走配对策略增强训练多样性
  • 在全局语义评估上超越现有SOTA方法
  • 适合关注自监督学习泛化性能的研究者

本文指出,从谱嵌入视角看,主流自监督学习方法VICReg可能存在次优性:过度依赖训练数据,导致对未见数据的泛化能力不足。为此提出SAG-VICReg(稳定且可泛化的VICReg),通过引入新训练技术,增强模型对数据全局语义的捕捉能力,并提升泛化性能。实验表明,SAG-VICReg在多种SOTA自监督学习基准上表现相当或更优,尤其在衡量全局语义理解的指标上显著领先,同时保持局部评价指标的竞争力。此外,我们提出一种无需标签的新独立嵌入评估指标,能有效反映无标签场景下的全局数据结构。

原文摘要 · Abstract (English)

In this paper, we argue that viewing VICReg-a popular self-supervised learning (SSL) method--through the lens of spectral embedding reveals a potential source of sub-optimality: it may struggle to generalize robustly to unseen data due to overreliance on the training data. This observation invites a closer look at how well this method achieves its goal of producing meaningful representations of images outside of the training set as well. Here, we investigate this issue and introduce SAG-VICReg (Stable and Generalizable VICReg), a method that builds on VICReg by incorporating new training techniques. These enhancements improve the model's ability to capture global semantics within the data and strengthen the generalization capabilities. Experiments demonstrate that SAG-VICReg effectively addresses the generalization challenge while matching or surpassing diverse state-of-the-art SSL baselines. Notably, our method exhibits superior performance on metrics designed to evaluate global semantic understanding, while simultaneously maintaining competitive results on local evaluation metrics. Furthermore, we propose a new standalone evaluation metric for embeddings that complements the standard evaluation methods and accounts for the global data structure without requiring labels--a key issue when tagged data is scarce or not available.

自监督学习泛化能力全局语义嵌入评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。