arXiv:2511.22125cs.CV2025-11

用通用属性锚点提升视频语言模型的泛化能力,防止微调时过拟合。

GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models

  • 引入预训练文本提示作为硬锚点,与软提示联合优化。
  • 在基础类到新类预测任务上显著优于现有方法。
  • 适合需要强泛化能力的视频理解场景。

视觉与文本软提示微调能有效提升视觉-语言模型在下游任务中的适应性。然而,在视频任务上微调会损害模型对未见类别的泛化能力。现有方法通过正则化手工设计提示与软提示之间的差异来缓解遗忘问题,但这也会削弱软提示的学习能力。为此,我们提出一种即插即用的耦合提示学习框架,以优化视觉-语言模型在视频任务中的泛化性能,核心思路是通过引入外部监督提示来缓解微调过程中的语义空间收缩。具体地,对于文本提示,我们引入来自其他数据集的预训练提示作为硬提示词元,将其与软提示词元拼接并通过可学习映射层耦合。这种竞争性提示策略防止了语义空间过度拟合于监督类别。此外,我们设计了一组无关视频集和负向提示作为通用属性锚点,以保持预训练语义空间中属性的通用相关性,从而维持模型的泛化能力。在视频任务上的实验表明,我们的方法在多个泛化基准上显著优于当前最优提示微调方法,尤其在基础类到新类预测任务上表现突出。

原文摘要 · Abstract (English)

Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the model's generalization ability to unseen classes. Existing methods attempt to mitigate this forgetting effect by regularizing the gap between hand-crafted prompts and soft prompts, but this also weakens the learning ability of soft prompts. To address this challenge, we propose a plug-and-play coupling prompt learning framework to optimize the generalization performance of V-L models in video tasks, with the core motivation of mitigating semantic space narrowing during fine-tuning by introducing an externally supervised prompt. Specifically, for textual prompts, we introduce pre-trained prompts from other datasets as hard prompt tokens. These are concatenated with soft prompt tokens and coupled via a learnable mapping layer. This competitive prompting approach prevents the semantic space from overfitting to supervised categories. In addition, we introduce a set of well-designed irrelevant video sets and negative prompts as generic attribute anchors to maintain the generic relevance of the attributes in the pre-trained semantic space, thus preserving the generalization ability. Experiments on video tasks demonstrate that our method significantly outperforms state-of-the-art prompt tuning approaches across generalization benchmarks, particularly on base-to-new class prediction.

提示微调视频理解泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。