arXiv:2504.08739cs.IRcs.HC2025-04

用手绘草图+语言智能,实现精准商品搜索推荐。

Enhancing Product Search Interfaces with Sketch-Guided Diffusion and Language Agents

  • 结合手绘草图与文本提示,通过扩散模型生成匹配图像。
  • 无需额外训练,准确率高,支持自然语言交互。
  • 适合电商、设计工具等需要直观搜索的场景。

扩散模型、Transformer 和语言代理的快速发展带来了新可能,但其在用户界面和商业应用中的潜力尚未充分挖掘。我们提出 Sketch-Search Agent,一种新颖框架,通过将多模态语言代理与自由手绘草图结合,作为扩散模型的控制信号,革新图像搜索体验。基于 T2I-Adapter,该框架融合草图与文本提示生成高质量查询图像,并通过 CLIP 图像编码器进行编码,以高效匹配图像库。相比现有方法,Sketch-Search Agent 无需额外训练、设置简单,在基于草图的图像检索与自然语言交互方面表现优异。多模态代理可动态保留用户偏好、排序结果并优化查询,提供个性化推荐。该交互式设计使用户可绘制草图并获得定制化产品建议,展示了扩散模型在以用户为中心的图像检索中的潜力。实验验证了其在提供相关产品搜索结果方面的高准确性。

原文摘要 · Abstract (English)

The rapid progress in diffusion models, transformers, and language agents has unlocked new possibilities, yet their potential in user interfaces and commercial applications remains underexplored. We present Sketch-Search Agent, a novel framework that transforms the image search experience by integrating a multimodal language agent with freehand sketches as control signals for diffusion models. Using the T2I-Adapter, Sketch-Search Agent combines sketches and text prompts to generate high-quality query images, encoded via a CLIP image encoder for efficient matching against an image corpus. Unlike existing methods, Sketch-Search Agent requires minimal setup, no additional training, and excels in sketch-based image retrieval and natural language interactions. The multimodal agent enhances user experience by dynamically retaining preferences, ranking results, and refining queries for personalized recommendations. This interactive design empowers users to create sketches and receive tailored product suggestions, showcasing the potential of diffusion models in user-centric image retrieval. Experiments confirm Sketch-Search Agent's high accuracy in delivering relevant product search results.

图像搜索扩散模型草图生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。