arXiv:2509.02024cs.CVcs.AI2025-09

用合成负样本提升视觉Transformer自监督学习效果

Unsupervised Training of Vision Transformers with Synthetic Negatives

  • 引入合成难负样本增强视觉Transformer表征能力
  • DeiT-S和Swin-T在多个任务上准确率显著提升
  • 适合关注自监督学习与模型泛化性的研究者

本文并未提出全新方法,而是关注自监督学习中被忽视的难负样本潜力。先前研究虽探索过合成难负样本,但极少应用于视觉Transformer。我们基于此观察,将合成难负样本融入视觉Transformer训练中,该简单有效的方法显著提升了学习表征的判别力。实验表明,该技术在DeiT-S和Swin-T架构上均取得性能提升。

原文摘要 · Abstract (English)

This paper does not introduce a novel method per se. Instead, we address the neglected potential of hard negative samples in self-supervised learning. Previous works explored synthetic hard negatives but rarely in the context of vision transformers. We build on this observation and integrate synthetic hard negatives to improve vision transformer representation learning. This simple yet effective technique notably improves the discriminative power of learned representations. Our experiments show performance improvements for both DeiT-S and Swin-T architectures.

视觉Transformer自监督学习负样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。