arXiv:2507.08064cs.MMcs.CV2025-07被引 17

轻量化多模态检索模型,高效且保持高精度。

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

  • 剪枝浅层网络并用深层特征做教师信号,减少参数量。
  • 按模态差异设计对比损失,提升训练效率。
  • 适合资源受限场景下的多模态检索应用。

随着多媒体内容增长,真实场景中对统一多模态检索(UMR)的需求日益增加。现有工作利用多模态大语言模型(MLLMs)解决此问题,但其庞大的参数量导致训练成本高、推理效率低。为此,我们提出PUMA:一种用于高效统一多模态检索的层剪枝语言模型,结合模态自适应学习。结构上,采用层剪枝自蒸馏,仅保留浅层网络,同时利用被移除深层的特征作为教师信号,降低参数量并保持表征能力。学习层面,引入模态自适应对比学习损失(MAC-Loss),根据目标模态将批次内负样本分为更难的同模态和较易的跨模态组,分别采用不同温度策略,提升学习效率。实验表明,该方法显著降低资源消耗,同时保持强性能。

原文摘要 · Abstract (English)

As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages multimodal large language models (MLLMs) to tackle this task. However, their large parameter size results in high training costs and low inference efficiency. To address this, we propose PUMA: a Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning. Our approach improves UMR from both structural and learning perspectives. (1) Structurally, we propose Layer-Pruned Self-Distillation, which prunes MLLMs by keeping only shallow layers while distilling features from dropped deep layers as teacher signals. This reduces parameters and preserves representation capability. (2) On the learning side, we introduce Modality-Adaptive Contrastive Learning Loss (MAC-Loss), which separates in-batch negatives into harder intra-modality and easier inter-modality groups based on the target modality, assigning different temperature strategies to enhance learning efficiency. Experiments show our method significantly reduces resource usage while maintaining strong performance.

多模态检索模型剪枝自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。