arXiv:2505.08819eess.IVcs.CV2025-05

改进稀疏分层图像掩码模式,提升自监督学习模型性能

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model

  • 设计新型网格状掩码模式,优化图像遮蔽方式
  • 在ImageNet上实现更高精度,超越原SparK模型
  • 适合研究自监督视觉表征与掩码策略的学者

掩码图像建模是自监督学习中最流行的训练目标之一。近期提出的SparK模型在同类方法中表现优异。本文针对该模型提出一种新的掩码模式,构建了网格掩码的SparK模型(Mesh Masked SparK)。研究系统评估了预训练阶段所用掩码模式对模型性能的影响,结果表明新掩码模式能有效提升表示能力,在ImageNet-1K上达到84.2%的分类准确率,优于原始SparK模型。

原文摘要 · Abstract (English)

Masked image modeling is one of the most poplular objectives of training. Recently, the SparK model has been proposed with superior performance among self-supervised learning models. This paper proposes a new mask pattern for this SparK model, proposing it as the Mesh Mask-ed SparK model. We report the effect of the mask pattern used for image masking in pre-training on performance.

自监督学习图像掩码SparK

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。