构建对话式图像检索数据集并提出端到端生成模型。
ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval
- 构建多轮图文交互的对话检索数据集
- 生成式模型在该数据集上表现优异
- 适合研究多模态交互与视觉问答的学者
本文研究开放域图像上的通用对话式图像检索任务,目标是通过人机交互对话来搜索图像。为此,我们构建了名为 ChatSearch 的数据集,包含每个目标图像对应的多轮多模态对话上下文查询,要求检索系统从数据库中精准定位对应图像。同时,我们提出一种名为 ChatSearcher 的生成式检索模型,可端到端处理交错的图文输入输出。该模型具备强大的多模态上下文推理能力,能利用世界知识生成视觉检索结果,在 ChatSearch 数据集上表现卓越,并在其他图像检索和视觉对话任务中也取得有竞争力的效果。我们期望该工作推动交互式多模态检索系统的进一步研究。数据集将开源至 https://github.com/joez17/ChatSearch。
原文摘要 · Abstract (English)
In this paper, we investigate the task of general conversational image retrieval on open-domain images. The objective is to search for images based on interactive conversations between humans and computers. To advance this task, we curate a dataset called ChatSearch. This dataset includes a multi-round multimodal conversational context query for each target image, thereby requiring the retrieval system to find the accurate image from database. Simultaneously, we propose a generative retrieval model named ChatSearcher, which is trained end-to-end to accept/produce interleaved image-text inputs/outputs. ChatSearcher exhibits strong capability in reasoning with multimodal context and can leverage world knowledge to yield visual retrieval results. It demonstrates superior performance on the ChatSearch dataset and also achieves competitive results on other image retrieval tasks and visual conversation tasks. We anticipate that this work will inspire further research on interactive multimodal retrieval systems. Our dataset will be available at https://github.com/joez17/ChatSearch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。