arXiv:2507.05677cs.CV2025-07被引 1

通过结构化提示增强视觉语言模型跨模态信息交互

Integrated Structural Prompt Learning for Vision-Language Models

  • 引入自结构与跨结构提示模块,建模提示与标记间的层级关系
  • 动态调整损失权重,提升对新类别的泛化能力
  • 在多个迁移任务中表现优于现有方法,尤其擅长处理新类别

提示学习方法显著提升了预训练视觉语言模型(如CLIP)在下游任务中的可迁移性。现有方法多采用人工模板或可学习向量提供文本或图像指令,但忽视了可学习提示与模态内及模态间词元之间的结构关系。同时,基类与新类性能平衡仍是难题。本文提出集成结构提示(ISP),增强视觉-语言分支间的信息表示交互。ISP引入自结构与跨结构提示模块,建模提示与冻结词元在模态内及跨模态的结构关系,实现高效信息传递并保持特征稳定性。此外,设计样本探测模块,根据样本难度动态调整损失系数,防止模型过拟合简单样本,提升对新类别的泛化能力。在三个主流设置(基类到新类泛化、跨数据集评估、域泛化)上的大量实验表明,ISP在性能上达到先进水平。

原文摘要 · Abstract (English)

Prompt learning methods have significantly extended the transferability of pre-trained Vision-Language Models (VLMs) like CLIP for various downstream tasks. These methods adopt handcraft templates or learnable vectors to provide text or image instructions in fine-tuning VLMs. However, most existing works ignore the structural relationships between learnable prompts and tokens within and between modalities. Moreover, balancing the performance of base and new classes remains a significant challenge. In this paper, we propose an Integrated Structural Prompt (ISP) for VLMs to enhance the interaction of information representations between the text and image branches. ISP introduces self-structural and cross-structural prompt modules to model the structural relationships between learnable prompts and frozen tokens within and across modalities. This enables efficient information transfer while preserving feature stability. Additionally, we propose a sample probing module that dynamically adjusts loss coefficients based on sample difficulty, preventing the mode from overfitting to simple samples and improving generalization ability to new classes. Extensive experiments on three widely used settings: base-to-new generalization, cross-dataset evaluation, and domain generalization demonstrate that the proposed ISP achieves competitive performance against state-of-the-art methods.

提示学习视觉语言模型跨模态交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。