arXiv:2409.05405cs.CVcs.AI2024-09综述被引 13

首篇系统综述多模态复合编辑与检索,梳理关键技术与未来方向。

A Survey of Multimodal Composite Editing and Retrieval

  • 系统归纳图文等多模态融合的编辑与检索方法
  • 覆盖应用、评测基准及实验结果,构建完整知识体系
  • 适合关注大模型多模态应用的研究者与工程师

现实世界中信息跨文本、图像、音频等多种模态,如何理解并利用这些数据以提升检索系统性能是研究重点。多模态复合检索通过整合文本、图像、音频等异构模态,提供更精准、个性化和上下文相关的结果。本综述深入探讨了图像-文本复合编辑、图像-文本复合检索及其他多模态复合检索任务,系统整理了应用场景、方法、基准数据集、实验结果与未来方向。在大模型时代,多模态学习备受关注,已有若干关于多模态学习与视觉-语言模型的综述发表于PAMI期刊。据我们所知,本综述是首个针对多模态复合检索领域的全面回顾,是对现有多模态融合综述的及时补充。为帮助读者快速追踪该领域进展,我们建立了项目主页:https://github.com/fuxianghuang1/Multimodal-Composite-Editing-and-Retrieval。

原文摘要 · Abstract (English)

In the real world, where information is abundant and diverse across different modalities, understanding and utilizing various data types to improve retrieval systems is a key focus of research. Multimodal composite retrieval integrates diverse modalities such as text, image and audio, etc. to provide more accurate, personalized, and contextually relevant results. To facilitate a deeper understanding of this promising direction, this survey explores multimodal composite editing and retrieval in depth, covering image-text composite editing, image-text composite retrieval, and other multimodal composite retrieval. In this survey, we systematically organize the application scenarios, methods, benchmarks, experiments, and future directions. Multimodal learning is a hot topic in large model era, and have also witnessed some surveys in multimodal learning and vision-language models with transformers published in the PAMI journal. To the best of our knowledge, this survey is the first comprehensive review of the literature on multimodal composite retrieval, which is a timely complement of multimodal fusion to existing reviews. To help readers' quickly track this field, we build the project page for this survey, which can be found at https://github.com/fuxianghuang1/Multimodal-Composite-Editing-and-Retrieval.

多模态检索综述大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。