arXiv:2505.08819eess.IVcs.CV2025-05
改进稀疏分层图像掩码模式,提升自监督学习模型性能
Thoughts on Objectives of Sparse and Hierarchical Masked Image Model
- 设计新型网格状掩码模式,优化图像遮蔽方式
- 在ImageNet上实现更高精度,超越原SparK模型
- 适合研究自监督视觉表征与掩码策略的学者
掩码图像建模是自监督学习中最流行的训练目标之一。近期提出的SparK模型在同类方法中表现优异。本文针对该模型提出一种新的掩码模式,构建了网格掩码的SparK模型(Mesh Masked SparK)。研究系统评估了预训练阶段所用掩码模式对模型性能的影响,结果表明新掩码模式能有效提升表示能力,在ImageNet-1K上达到84.2%的分类准确率,优于原始SparK模型。
原文摘要 · Abstract (English)
Masked image modeling is one of the most poplular objectives of training. Recently, the SparK model has been proposed with superior performance among self-supervised learning models. This paper proposes a new mask pattern for this SparK model, proposing it as the Mesh Mask-ed SparK model. We report the effect of the mask pattern used for image masking in pre-training on performance.
自监督学习图像掩码SparK
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。