提出轻量级图像引导检索数据集,兼顾视觉与文本查询性能。
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
- 设计可选文本的统一检索框架,融合视觉与组合查询能力。
- 在16,474个训练样本上实现34.8 mAP@10和75.7 mAP@200的高精度。
- 适合关注高效跨模态检索的开发者与研究者使用。
图像引导检索中的可选文本(IGROT)任务统一了纯视觉检索与组合式检索。尽管在谷歌图片、必应等应用中具有重要价值,但受限于缺乏易用的基准数据集及平衡各子任务表现的方法。现有大规模数据集如MagicLens虽全面但计算成本高,而多数模型偏向视觉或组合查询。本文提出FIGROTD,一个轻量且高质量的IGROT数据集,包含16,474个训练三元组和1,262个测试三元组,覆盖CIR、SBIR和CSTBIR任务。为减少冗余,提出基于方差的特征掩码(VaGFeM),按方差统计选择性增强判别维度,并采用双损失设计(InfoNCE + Triplet)提升组合推理能力。在FIGROTD上训练的模型,在九个基准上表现优异,达34.8 mAP@10(CIRCO)和75.7 mAP@200(Sketchy),优于更强基线,即使训练样本更少。
原文摘要 · Abstract (English)
Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Image and Bing, progress has been limited by the lack of an accessible benchmark and methods that balance performance across subtasks. Large-scale datasets such as MagicLens are comprehensive but computationally prohibitive, while existing models often favor either visual or compositional queries. We introduce FIGROTD, a lightweight yet high-quality IGROT dataset with 16,474 training triplets and 1,262 test triplets across CIR, SBIR, and CSTBIR. To reduce redundancy, we propose the Variance Guided Feature Mask (VaGFeM), which selectively enhances discriminative dimensions based on variance statistics. We further adopt a dual-loss design (InfoNCE + Triplet) to improve compositional reasoning. Trained on FIGROTD, VaGFeM achieves competitive results on nine benchmarks, reaching 34.8 mAP@10 on CIRCO and 75.7 mAP@200 on Sketchy, outperforming stronger baselines despite fewer triplets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。