arXiv:2509.06566cs.CV2025-09中稿 · BMVC2025被引 1

针对手绘草图的图像检索,提出更鲁棒的训练方法。

Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval

  • 设计抗草图噪声的训练目标,提升对真实草图变异性的适应能力。
  • 在FS-COCO和SketchyCOCO上达到当前最优性能。
  • 适合关注跨模态检索训练策略与评估标准的研究者。

场景级草图图像检索旨在找到与自由手绘草图整体语义和空间布局匹配的自然图像。不同于以往聚焦于模型结构改进的工作,本文强调真实草图固有的模糊性与噪声问题,提出一种显式设计以抵御草图变异性。通过合理组合预训练、编码器架构与损失函数,无需引入额外复杂性即可实现最先进性能。在具有挑战性的FS-COCO和广泛使用的SketchyCOCO数据集上的大量实验验证了该方法的有效性,凸显了训练设计在跨模态检索任务中的关键作用,以及改进场景级草图图像检索评估范式的重要性。

原文摘要 · Abstract (English)

The goal of Scene-level Sketch-Based Image Retrieval is to retrieve natural images matching the overall semantics and spatial layout of a free-hand sketch. Unlike prior work focused on architectural augmentations of retrieval models, we emphasize the inherent ambiguity and noise present in real-world sketches. This insight motivates a training objective that is explicitly designed to be robust to sketch variability. We show that with an appropriate combination of pre-training, encoder architecture, and loss formulation, it is possible to achieve state-of-the-art performance without the introduction of additional complexity. Extensive experiments on a challenging FS-COCO and widely-used SketchyCOCO datasets confirm the effectiveness of our approach and underline the critical role of training design in cross-modal retrieval tasks, as well as the need to improve the evaluation scenarios of scene-level SBIR.

草图检索跨模态训练设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。