arXiv:2505.13306cs.CVcs.IR2025-05被引 6

用高斯混合模型捕捉少样本跨模态数据的多峰分布,提升检索准确率。

GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

  • 用GMM建模跨模态数据的多峰分布,避免单峰假设带来的偏差。
  • 通过多正样本对比学习,增强特征表达的全面性与鲁棒性。
  • 约束图像与文本特征间的相对距离,缓解语义鸿沟,适合少样本场景。

少样本跨模态检索旨在仅用少量训练样本学习跨模态表示,使模型能处理推理时的未见类别。与传统跨模态检索不同,少样本任务中模态间数据分布稀疏,现有方法常无法有效建模数据的多峰特性,导致潜在语义空间中存在两类偏差:模态内偏差(稀疏样本难以捕捉类内多样性)和模态间偏差(图像与文本分布错位加剧语义鸿沟),从而影响检索精度。为此,本文提出一种新方法GCRDP,利用高斯混合模型(GMM)有效捕捉数据的复杂多峰分布,并引入多正样本对比学习机制实现全面特征建模。同时,设计了一种新的跨模态语义对齐策略,通过约束图像与文本特征分布间的相对距离,提升跨模态表示准确性。在四个基准数据集上的大量实验表明,该方法优于六种最先进的方法。

原文摘要 · Abstract (English)

Few-shot cross-modal retrieval focuses on learning cross-modal representations with limited training samples, enabling the model to handle unseen classes during inference. Unlike traditional cross-modal retrieval tasks, which assume that both training and testing data share the same class distribution, few-shot retrieval involves data with sparse representations across modalities. Existing methods often fail to adequately model the multi-peak distribution of few-shot cross-modal data, resulting in two main biases in the latent semantic space: intra-modal bias, where sparse samples fail to capture intra-class diversity, and inter-modal bias, where misalignments between image and text distributions exacerbate the semantic gap. These biases hinder retrieval accuracy. To address these issues, we propose a novel method, GCRDP, for few-shot cross-modal retrieval. This approach effectively captures the complex multi-peak distribution of data using a Gaussian Mixture Model (GMM) and incorporates a multi-positive sample contrastive learning mechanism for comprehensive feature modeling. Additionally, we introduce a new strategy for cross-modal semantic alignment, which constrains the relative distances between image and text feature distributions, thereby improving the accuracy of cross-modal representations. We validate our approach through extensive experiments on four benchmark datasets, demonstrating superior performance over six state-of-the-art methods.

少样本学习跨模态检索GMM对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。