arXiv:2503.15060cs.CVcs.AI2025-03

提出Sorcen框架,用生成式正样本统一表征学习与图像合成。

Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis

  • 通过生成语义令牌中的回声样本来构建对比正例,无需额外数据增强。
  • 在ImageNet-1k上比现有统一自监督方法线性探测提升0.4%,生成质量提高1.48 FID。
  • 仅用预计算令牌,训练效率提升60.8%,适合高效统一模型研究者。

表征学习与生成建模虽旨在理解视觉数据,但二者统一仍待探索。现有统一自监督学习(SSL)方法依赖语义令牌重建,需外部分词器,带来显著开销。本文提出Sorcen,一种新型统一SSL框架,采用协同对比-重构目标。其对比目标“回声对比”(Echo Contrast)利用Sorcen的生成能力,在语义令牌空间生成回声样本作为对比正例,无需额外图像裁剪或增强。Sorcen仅处理预计算令牌,避免训练时在线令牌转换,大幅降低计算开销。ImageNet-1k上的大量实验表明,Sorcen在线性探测、无条件图像生成、少样本学习和迁移学习上分别优于先前统一自监督方法0.4%、1.48 FID、1.76%和1.53%,且效率提升60.8%。此外,其线性探测性能超越单裁剪掩码建模(MIM)最先进水平,并在无条件图像生成上达到最先进性能,彰显统一自监督模型的重大突破。

原文摘要 · Abstract (English)

While representation learning and generative modeling seek to understand visual data, unifying both domains remains unexplored. Recent Unified Self-Supervised Learning (SSL) methods have started to bridge the gap between both paradigms. However, they rely solely on semantic token reconstruction, which requires an external tokenizer during training -- introducing a significant overhead. In this work, we introduce Sorcen, a novel unified SSL framework, incorporating a synergic Contrastive-Reconstruction objective. Our Contrastive objective, "Echo Contrast", leverages the generative capabilities of Sorcen, eliminating the need for additional image crops or augmentations during training. Sorcen "generates" an echo sample in the semantic token space, forming the contrastive positive pair. Sorcen operates exclusively on precomputed tokens, eliminating the need for an online token transformation during training, thereby significantly reducing computational overhead. Extensive experiments on ImageNet-1k demonstrate that Sorcen outperforms the previous Unified SSL SoTA by 0.4%, 1.48 FID, 1.76%, and 1.53% on linear probing, unconditional image generation, few-shot learning, and transfer learning, respectively, while being 60.8% more efficient. Additionally, Sorcen surpasses previous single-crop MIM SoTA in linear probing and achieves SoTA performance in unconditional image generation, highlighting significant improvements and breakthroughs in Unified SSL models.

统一学习生成建模自监督高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。