针对产品设计检索中结构匹配需求,构建了多模态评估数据集
Sketch2Inspire: Structure-Sensitive Evaluation for Product Retrieval

- 基于边缘草图与文本查询融合,分离类别与结构检索任务
- 多模态融合在结构敏感评估中表现最优(nDCG=0.7015)
- 提供人工标注基准,适合结构感知检索系统开发
早期产品设计检索不仅需要类别识别,还需匹配语义意图和粗略结构特征。现有图像资源和通用图文检索基准难以区分类别检索与类内结构匹配。本文构建了Sketch2Inspire数据集,源自Amazon Berkeley Objects的精选子集,包含对齐的文本查询、基于边缘的草图代理查询及文本-草图融合查询。该资源分离了广义类别检索与结构敏感的类内检索,并引入人工评分参考协议进行校准。我们评估了一个基于预训练CLIP族编码器的轻量级参考系统,对比纯文本检索、纯草图检索、加权后期融合及无需更新权重的文本优先重排序。在广义相关性下,后期融合得分最高(nDCG = 0.9962);在自动结构敏感相关性下,其表现依然最优(nDCG = 0.7015),显著优于纯文本检索(nDCG = 0.5912)。人工评分结果显示,后期融合在nDCG@10上最高(0.9133),纯文本次之(0.9030)。结果表明,多模态输入的检索增益取决于相关性定义方式。Sketch2Inspire因此成为诊断模态贡献、支持结构感知检索协议发展的关键资源。
原文摘要 · Abstract (English)
Early-stage product design retrieval often requires more than category recognition: designers may need reference examples that match both a short semantic intent and a rough structural cue. Existing product-image resources and generic image--text retrieval benchmarks rarely separate category retrieval from within-category structural fit. We present Sketch2Inspire, built from a curated subset of Amazon Berkeley Objects with aligned text queries, edge-based sketch-proxy queries, and fused text--sketch queries. The resource separates broad category-level retrieval from structure-sensitive within-category retrieval and includes a human-graded reference protocol for calibration. We evaluate a lightweight reference system based on pretrained CLIP-family encoders, comparing text-only retrieval, sketch-only retrieval, weighted late fusion, and text-first reranking without updating model weights. Under broad relevance, late fusion obtains the highest score (nDCG = 0.9962). Under automatic structure-sensitive relevance, late fusion again obtains the highest score (nDCG = 0.7015), exceeding text-only retrieval (nDCG = 0.5912). In the human-graded results, late fusion obtains the highest nDCG@10 (0.9133), while text-only retrieval ranks second (0.9030). These results show that the retrieval gain from multimodal input depends on how relevance is defined. Sketch2Inspire therefore provides a diagnostic resource for evaluating modality contribution and supports the development of structure-aware product-retrieval protocols with independent human annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。