让不会画画的人也能用手绘+文字找难名难画的物体。
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
- 用草图定位物体,文本描述属性或动作,联合检索。
- 在200万条混合查询上,超越纯文本/纯草图方法。
- 适合非母语者、儿童或复杂场景搜索,支持跨模态交互。
非母语者常因词汇量有限难以命名具体物体(如澳大利亚的袋鼬),即便能想象其形态。此外,用户可能希望搜索难以绘制但易于描述的复杂互动(如袋鼬挖土)。针对这类常见但复杂的查询需求,现有文本或草图图像检索方法均不适用。为此,我们构建了包含约200万条查询和10.8万张自然场景图像的全新数据集CSTBIR(复合草图+文本图像检索)。并提出基于预训练多模态变换器的基准模型STNET:利用手绘草图定位图像中相关物体,并结合文本与图像编码进行检索。除对比学习外,设计多种训练目标提升性能。大量实验表明,该方法在纯文本、纯草图及复合查询任务上均优于当前最优模型。数据集与代码已公开。
原文摘要 · Abstract (English)
Non-native speakers with limited vocabulary often struggle to name specific objects despite being able to visualize them, e.g., people outside Australia searching for numbats. Further, users may want to search for such elusive objects with difficult-to-sketch interactions, e.g., numbat digging in the ground. In such common but complex situations, users desire a search interface that accepts composite multimodal queries comprising hand-drawn sketches of difficult-to-name but easy-to-draw objects and text describing difficult-to-sketch but easy-to-verbalize object attributes or interaction with the scene. This novel problem statement distinctly differs from the previously well-researched TBIR (text-based image retrieval) and SBIR (sketch-based image retrieval) problems. To study this under-explored task, we curate a dataset, CSTBIR (Composite Sketch+Text Based Image Retrieval), consisting of approx. 2M queries and 108K natural scene images. Further, as a solution to this problem, we propose a pretrained multimodal transformer-based baseline, STNET (Sketch+Text Network), that uses a hand-drawn sketch to localize relevant objects in the natural scene image, and encodes the text and image to perform image retrieval. In addition to contrastive learning, we propose multiple training objectives that improve the performance of our model. Extensive experiments show that our proposed method outperforms several state-of-the-art retrieval methods for text-only, sketch-only, and composite query modalities. We make the dataset and code available at our project website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。