arXiv:2606.30393cs.CV2026-06

提出首个关注主体的干扰物定位基准,帮图像编辑自动识别该删哪些

SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization

论文配图:SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization
图 1 · 摘自论文原文
  • 构建主体感知干扰物定位任务,明确哪些对象该保留、哪些该删
  • 包含1800个真实场景案例,14617个标注候选,含1938个难负样本
  • 揭示当前视觉语言模型过度删除问题,为多模态推理提供诊断工具

照片中常存在干扰主体注意力的视觉干扰物,影响构图效果。尽管现代编辑工具可快速移除物体,但判断移除目标仍依赖人工。现有显著性模型和开放词汇检测器缺乏主体意识,无法适应用户意图变化;且不考虑上下文的盲目删除可能破坏场景语义一致性(如保留人物却移除其坐的椅子)。为此,本文首次形式化了主体感知干扰物定位任务,旨在识别干扰物的同时保留构图必需对象。提出 extsc{SADL},首个真实世界基准,涵盖1000张照片中的1800个主体感知案例,共14617个标注候选,其中包含1938个难负样本用于压力测试排除校准。评估七种专有及开源视觉语言模型(VLMs),采用先分类后过滤的流水线,基于五类包含因素与三类上下文排除规则进行评测。分析表明,VLMs 在识别干扰物方面表现良好,但过度应用排除策略,导致大规模系统性抑制真实干扰物。 extsc{SADL} 揭示此关键瓶颈,为多模态系统中的主体条件推理提供基础诊断工具。

原文摘要 · Abstract (English)

Photographs frequently contain \emph{visual distractors} besides foregrounds and backgrounds of the intended subject, competing for attention and weakening composition. While modern editing tools streamline object removal, identifying which objects to remove remains a mostly manual process. Existing saliency models and open-vocabulary detectors operate without subject awareness, failing to adapt to shifting user intent. Furthermore, context-agnostic removal may disrupt the scene's semantic coherence (e.g., keep the person but remove the chair they are sitting on). To address these limitations, we formalize the task of subject-aware distractor localization, which identifies distractors while retaining compositionally essential objects. This paper introduces \textsc{SADL}, the first real-world benchmark for this task, comprising 1,800 subject-aware cases across 1,000 photographs to enable systematic evaluation and facilitate future research. In total, there are 14,617 annotated candidates, including a robust set of 1,938 hard negatives to stress-test exclusion calibration. We evaluate seven proprietary and open-weight Vision-Language Models (VLMs) on a sequential pipeline of distractor classification followed by exclusion filtering, structured around five inclusion factors and three contextual exclusion rules. Our analysis reveals that VLMs are highly capable of identifying distractors, but then over-apply exclusion, which systematically suppresses true distractors at scale. By exposing this critical bottleneck, \textsc{SADL} provides a foundational diagnostic tool to advance subject-conditioned reasoning in multimodal systems.

图像编辑视觉语言模型干扰物定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。