arXiv:2503.01334cs.IRcs.CV2025-03综述被引 4

通过图文组合查询,实现更精准的多模态内容检索。

Composed Multi-modal Retrieval: A Survey of Approaches and Applications

  • 结合视觉参考与文本修改进行检索,提升查询灵活性。
  • 涵盖监督、零样本和半监督三种学习范式,支持不同场景需求。
  • 适用于电商、安防等实际应用,推动智能搜索发展。

多模态数据的快速增长催生了超越单模态与跨模态检索的新范式。组合式多模态检索(Composed Multi-modal Retrieval, CMR)作为下一代关键技术,允许用户通过引入参考图像或视频并辅以文本修改来查询内容,实现前所未有的灵活性与精确性。本文全面综述了CMR的基础挑战、技术进展与应用场景。根据学习范式,CMR被分为监督、零样本与半监督三类。文中讨论了监督学习中的数据构建、模型架构与损失优化,零样本学习中的转换框架与线性融合机制,以及半监督方法如何利用生成伪三元组并应对数据噪声。此外,系统梳理了CMR在电商、社交媒体、搜索引擎、公共安全等领域的广泛应用,深入分析了七个高影响力场景,涵盖基准数据集与性能评估。最后,提出若干新兴研究方向,以期激发在未探索领域中的创新。相关工作清单见:https://github.com/kkzhang95/Awesome-Composed-Multi-modal-Retrieval

原文摘要 · Abstract (English)

The burgeoning volume of multi-modal data necessitates advanced retrieval paradigms beyond unimodal and cross-modal approaches. Composed Multi-modal Retrieval (CMR) emerges as a pivotal next-generation technology, enabling users to query images or videos by integrating a reference visual input with textual modifications, thereby achieving unprecedented flexibility and precision. This paper provides a comprehensive survey of CMR, covering its fundamental challenges, technical advancements, and applications. CMR is categorized into supervised, zero-shot, and semi-supervised learning paradigms. We discuss key research directions, including data construction, model architecture, and loss optimization in supervised CMR, as well as transformation frameworks and linear integration in zero-shot CMR, and semi-supervised CMR that leverages generated pseudo-triplets while addressing data noise/uncertainty. Additionally, we extensively survey the diverse application landscape of CMR, highlighting its transformative potential in e-commerce, social media, search engines, public security, etc. Seven high impact application scenarios are explored in detail with benchmark data sets and performance analysis. Finally, we further provide new potential research directions with the hope of inspiring exploration in other yet-to-be-explored fields. A curated list of works is available at: https://github.com/kkzhang95/Awesome-Composed-Multi-modal-Retrieval

多模态检索组合查询智能搜索应用综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。