arXiv:2604.00172cs.CV2026-04

提出方法抑制图像建模中的非语义噪声,提升零样本性能

Suppressing Non-Semantic Noise in Masked Image Modeling Representations

  • 用PCA分析真实与合成的非语义图像,构建语义不变性评分
  • 引入SOAP方法,通过线性头直接抑制块表示中的非语义信息
  • 无需训练、可适配任意模型,显著提升多种MIM模型的零样本表现

掩码图像建模(Masked Image Modeling, MIM)已成为主流的自监督视觉范式。本文发现,MIM目标会导致学习到的表征保留非语义信息,从而损害推理阶段性能。我们基于真实与合成非语义图像的主成分分析(PCA),提出一种模型无关的语义不变性评分。据此,提出一种简单方法——语义正交伪影投影(Semantically Orthogonal Artifact Projection, SOAP),可直接抑制块表示中的非语义信息,实现对多种MIM基模型一致的零样本性能提升。SOAP为后处理抑制方法,无需训练,可作为单一线性头附加至任何模型。

原文摘要 · Abstract (English)

Masked Image Modeling (MIM) has become a ubiquitous self-supervised vision paradigm. In this work, we show that MIM objectives cause the learned representations to retain non-semantic information, which ultimately hurts performance during inference. We introduce a model-agnostic score for semantic invariance using Principal Component Analysis (PCA) on real and synthetic non-semantic images. Based on this score, we propose a simple method, Semantically Orthogonal Artifact Projection (SOAP), to directly suppress non-semantic information in patch representations, leading to consistent improvements in zero-shot performance across various MIM-based models. SOAP is a post-hoc suppression method, requires zero training, and can be attached to any model as a single linear head.

自监督学习图像建模表征优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。