用结构化语义标识提升跨模态生成检索的准确率
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
- 用自然语言结构化标识增强语义对齐
- 文本到图像检索召回率提升8.65点
- 首个统一处理双向检索的生成框架
跨模态检索(CMR)是多媒体研究中的基础任务,旨在跨不同模态检索语义相关目标。传统方法通过嵌入相似度匹配文本与图像,而预训练生成模型的进展使生成式检索成为有前景的替代方案。该范式为每个目标分配唯一标识符,并利用生成模型直接预测与输入查询对应的标识符,无需显式索引。尽管潜力巨大,现有生成式CMR方法在标识构建和生成过程中仍存在语义信息不足的问题。为此,我们提出一种新型统一的语义增强型生成式跨模态检索框架(SemCORE),旨在释放生成式跨模态检索中的语义理解能力。具体而言,我们首先构建结构化自然语言标识符(SID),有效对齐目标标识与优化于自然语言理解与生成的生成模型;进一步引入生成式语义验证(GSV)策略,实现细粒度目标区分。此外,据我们所知,SemCORE是首个在生成式跨模态检索中同时考虑文本到图像和图像到文本检索任务的框架。大量实验表明,本框架显著优于当前最优的生成式跨模态检索方法。特别地,其在基准数据集上实现了显著提升,文本到图像检索的Recall@1平均提高8.65点。
原文摘要 · Abstract (English)
Cross-modal retrieval (CMR) is a fundamental task in multimedia research, focused on retrieving semantically relevant targets across different modalities. While traditional CMR methods match text and image via embedding-based similarity calculations, recent advancements in pre-trained generative models have established generative retrieval as a promising alternative. This paradigm assigns each target a unique identifier and leverages a generative model to directly predict identifiers corresponding to input queries without explicit indexing. Despite its great potential, current generative CMR approaches still face semantic information insufficiency in both identifier construction and generation processes. To address these limitations, we propose a novel unified Semantic-enhanced generative Cross-mOdal REtrieval framework (SemCORE), designed to unleash the semantic understanding capabilities in generative cross-modal retrieval task. Specifically, we first construct a Structured natural language IDentifier (SID) that effectively aligns target identifiers with generative models optimized for natural language comprehension and generation. Furthermore, we introduce a Generative Semantic Verification (GSV) strategy enabling fine-grained target discrimination. Additionally, to the best of our knowledge, SemCORE is the first framework to simultaneously consider both text-to-image and image-to-text retrieval tasks within generative cross-modal retrieval. Extensive experiments demonstrate that our framework outperforms state-of-the-art generative cross-modal retrieval methods. Notably, SemCORE achieves substantial improvements across benchmark datasets, with an average increase of 8.65 points in Recall@1 for text-to-image retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。