arXiv:2604.15735cs.CVcs.AI2026-04

用草图结构+文本描述提升细粒度图像检索效果

Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval

论文配图:Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval
图 1 · 摘自论文原文
  • 融合草图轮廓与文本属性,互补信息增强表征
  • 在自建数据集上优于现有方法,显著提升检索准确率
  • 适合需要跨模态图文检索的研究者与应用开发

通过手绘草图或文本描述进行细粒度图像检索仍面临模态鸿沟的挑战。草图能捕捉复杂结构轮廓,但缺乏颜色与纹理;文本可提供丰富的颜色和纹理信息,却丢失空间布局。针对二者互补性,本文提出草图与文本驱动的图像检索框架(STBIR)。通过融合文本中的颜色、纹理线索与草图的结构轮廓,显著提升检索性能。首先,设计基于课程学习的鲁棒性增强模块,提升模型对不同质量查询的适应能力;其次,引入基于类别知识的特征空间优化模块,增强模型表达能力;最后,构建多阶段跨模态特征对齐机制,有效缓解跨模态对齐难题。此外,我们构建了细粒度STBIR基准数据集,用于严格验证框架有效性,并为后续研究提供数据支持。大量实验表明,所提STBIR框架显著优于现有先进方法。

原文摘要 · Abstract (English)

Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex structural contours, they lack color and texture, which text effectively provides despite omitting spatial contours. Motivated by the complementary nature of these modalities, we propose the Sketch and Text Based Image Retrieval (STBIR) framework. By synergizing the rich color and texture cues from text with the structural outlines provided by sketches, STBIR achieves superior fine-grained retrieval performance. First, a curriculum learning driven robustness enhancement module is proposed to enhance the model's robustness when handling queries of varying quality. Second, we introduce a category-knowledge-based feature space optimization module, thereby significantly boosting the model's representational power. Finally, we design a multi-stage cross-modal feature alignment mechanism to effectively mitigate the challenges of cross modal feature alignment. Furthermore, we curate the fine-grained STBIR benchmark dataset to rigorously validate the efficacy of our proposed framework and to provide data support as a reference for subsequent related research. Extensive experiments demonstrate that the proposed STBIR framework significantly outperforms state of the art methods.

细粒度检索跨模态草图生成特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。