arXiv:2604.03980cs.CVcs.AI2026-04

用视觉特征的统计结构增强语言提示,提升模型跨域适应能力

Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics

论文配图:Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
图 1 · 摘自论文原文
  • 引入格拉姆矩阵捕捉视觉特征的二阶统计特性,补充一阶空间信息
  • 在多个基准上优于传统方法,显著提升跨域任务表现
  • 适合需要强鲁棒性的视觉语言模型微调场景

参数高效提示学习已成为将视觉-语言模型(VLMs)适配到下游任务的主流方法。现有方法主要关注将文本提示与一阶视觉特征(即空间特征图)对齐。虽然在细粒度语义区分上有效,但仅依赖一阶信息不足以实现稳健适配,因为这些空间耦合的特征极易受领域偏移和局部噪声影响。本文提出基于二阶统计的格拉姆锚定提示学习(GAPL),通过引入额外的二阶统计流(以格拉姆矩阵实现),增强标准的一阶空间交互。通过将提示锚定于这些二阶先验,该方法使语言表示能动态适应不同领域间的统计分布变化。大量实验表明,二阶特征具有显著有效性,GAPL在多个基准上展现出卓越性能。

原文摘要 · Abstract (English)

Parameter-efficient prompt learning has become the de facto standard for adapting Vision-Language Models (VLMs) to downstream tasks. Existing approaches predominantly focus on aligning text prompts with first-order visual features (i.e., spatial feature maps). While effective for fine-grained semantic discrimination, we argue that relying solely on first-order information is insufficient for robust adaptation, as these spatially entangled features are highly susceptible to domain shifts and local noise. In this work, we propose \textbf{Gram-Anchored Prompt Learning (GAPL)} for Vision-Language Models via Second-Order Statistics, a framework that synergizes local semantic alignment with global structural consistency. Methodologically, we introduce an additional second-order statistical stream via \textbf{Gram matrices} that augments the standard first-order spatial interaction. By anchoring prompts to these second-order priors, our approach enables language representations to dynamically adapt to statistical distribution shifts across diverse domains. Extensive experiments indicate the effectiveness of the second-order features, and show compelling performances of GAPL on various benchmarks.

视觉语言模型提示学习二阶统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。