arXiv:2607.09164cs.CV2026-07

用可学习掩码直接优化保留模型行为的最小像素区域。

What Pixels Are Enough? SEAMS: Sufficiency Saliency via MSE-Preservation Soft-Masks

论文配图:What Pixels Are Enough? SEAMS: Sufficiency Saliency via MSE-Preservation Soft-Masks
图 1 · 摘自论文原文
  • 通过软掩码与可学习预算,直接优化保留指定输出
  • 生成紧凑可解释掩码,在插入删除任务上表现优异
  • 适用于不同架构与目标,揭示模型依赖的差异证据

显著性图最有价值时,是能识别出足以维持模型行为的图像区域。我们提出SEAMS,一种基于充分性的显著性方法,通过保全目标直接优化软掩码。给定冻结的可微模型输出(如类别概率、CLS嵌入或标记表示),SEAMS搜索能保留选定输出的紧凑掩码。该方法基于简单的优化框架:软掩码、可学习预算,以及完全由查询图像生成的三路图像合成。因此无需额外干扰数据集、特定架构的归因机制或可微top-k近似。在冻结的ViT-S/16和ConvNeXt模型上的实验表明,仅改变保留目标,同一优化流程即可生成物体级、类别条件和标记级解释。所得掩码紧凑、可解释、对随机初始化稳定,且在插入与删除基准测试中表现良好。结果还表明,不同架构通常依赖不同的充分证据,却能达到相似的保全精度,凸显视觉解释的架构依赖性。

原文摘要 · Abstract (English)

Saliency maps are most useful when they identify the image regions that are sufficient to preserve a model's behaviour. We introduce SEAMS, a sufficiency-based saliency method that directly optimises a soft mask using a preservation objective. Given a frozen differentiable model output, such as a class probability, CLS embedding, or token representation, SEAMS searches for a compact mask that preserves the selected output. The approach relies on a simple optimisation framework based on soft masks, a learnable budget, and a three-way image composite generated entirely from the query image. As a result, it requires no auxiliary distractor dataset, architecture-specific attribution mechanism, or differentiable top-k relaxation. Experiments with frozen ViT-S/16 and ConvNeXt models show that the same optimisation pipeline can generate object-level, class-conditioned, and token-level explanations by changing only the preserved target. The resulting masks are compact, interpretable, stable across random initialisations, and competitive on insertion and deletion benchmarks. Our results also indicate that different architectures often rely on different sufficient evidence while achieving similar preservation fidelity, highlighting the architecture-dependent nature of visual explanations.

显著性分析视觉解释掩码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。