arXiv:2511.12480cs.CVcs.AI2025-11AAAI被引 1

让被遮盖的图像区域重获价值,提升模型细节捕捉能力

MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning

论文配图:MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning
图 1 · 摘自论文原文
  • 将遮蔽区域视为辅助信息,通过重学机制联合学习可见与遮蔽内容
  • 在多个基准上实现一致性能提升,尤其增强细粒度任务表现
  • 适用于任意主干网络,适合注重细节建模的视觉任务

监督学习中的传统图像遮蔽存在两个关键问题:(i) 被丢弃的像素未被充分利用,导致有价值的情境信息丢失;(ii) 遮蔽可能移除小而关键的特征,尤其在细粒度任务中。相比之下,掩码图像建模(MIM)表明,即使部分输入也能重建遮蔽区域,揭示了不完整数据与原图之间存在强情境一致性,凸显了遮蔽区域蕴含的语义多样性潜力。受此启发,我们重新思考图像遮蔽策略,提出将遮蔽内容视为辅助知识而非忽略对象。基于此,我们提出 MaskAnyNet,结合遮蔽与重学机制,同时利用可见与遮蔽信息。该方法可轻松扩展至任意具有附加分支的模型,实现对重构遮蔽区域的联合学习。该方法通过重用遮蔽内容提升特征语义多样性并保留细粒度细节。在CNN与Transformer主干网络上的实验显示,在多个基准上均取得稳定增益。进一步分析证实,该方法通过复用遮蔽内容增强了语义多样性。

原文摘要 · Abstract (English)

In supervised learning, traditional image masking faces two key issues: (i) discarded pixels are underutilized, leading to a loss of valuable contextual information; (ii) masking may remove small or critical features, especially in fine-grained tasks. In contrast, masked image modeling (MIM) has demonstrated that masked regions can be reconstructed from partial input, revealing that even incomplete data can exhibit strong contextual consistency with the original image. This highlights the potential of masked regions as sources of semantic diversity. Motivated by this, we revisit the image masking approach, proposing to treat masked content as auxiliary knowledge rather than ignored. Based on this, we propose MaskAnyNet, which combines masking with a relearning mechanism to exploit both visible and masked information. It can be easily extended to any model with an additional branch to jointly learn from the recomposed masked region. This approach leverages the semantic diversity of the masked regions to enrich features and preserve fine-grained details. Experiments on CNN and Transformer backbones show consistent gains across multiple benchmarks. Further analysis confirms that the proposed method improves semantic diversity through the reuse of masked content.

图像遮蔽语义多样性细粒度识别特征增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。