arXiv:2504.01952cs.CV2025-04

让AI根据文字指令精准定位多图间的细微差异。

Image Difference Grounding with Natural Language

  • 基于自然语言指令,精确定位图像对中的视觉差异。
  • 提出新数据集DiffGround,含多样变化的图像对与细粒度指令。
  • 适合智能监控、医学影像对比等需精准差异检测场景。

视觉定位(VG)通常聚焦于用自然语言定位单张图像中的感兴趣区域,但现有方法难以处理多图场景。在自动监控等真实场景中,发现细微但有意义的视觉差异至关重要。此前图像差异理解(IDU)研究或仅检测所有变化区域而无文本引导,或仅提供粗粒度描述。为此,我们提出图像差异定位(IDG)任务,旨在根据用户指令精准定位视觉差异。构建了大规模高质量的DiffGround数据集,包含具有多样化视觉变化的图像对及细粒度差异查询指令。同时提出基线模型DiffTracker,通过特征差异增强与共性抑制实现精确差异定位。在DiffGround上的实验表明,该数据集对实现更细粒度的IDU至关重要。为促进后续研究,将公开DiffGround数据集与DiffTracker模型。

原文摘要 · Abstract (English)

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretations. This limits their applicability in real-world scenarios like automatic surveillance, where detecting subtle but meaningful visual differences across multiple images is crucial. Besides, previous work on image difference understanding (IDU) has either focused on detecting all change regions without cross-modal text guidance, or on providing coarse-grained descriptions of differences. Therefore, to push towards finer-grained vision-language perception, we propose Image Difference Grounding (IDG), a task designed to precisely localize visual differences based on user instructions. We introduce DiffGround, a large-scale and high-quality dataset for IDG, containing image pairs with diverse visual variations along with instructions querying fine-grained differences. Besides, we present a baseline model for IDG, DiffTracker, which effectively integrates feature differential enhancement and common suppression to precisely locate differences. Experiments on the DiffGround dataset highlight the importance of our IDG dataset in enabling finer-grained IDU. To foster future research, both DiffGround data and DiffTracker model will be publicly released.

视觉定位多图对比语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。