arXiv:2504.14123cs.AIcs.CL2025-04被引 3

用贝叶斯原理改进视觉语言模型的提示学习,提升泛化能力

Bayesian Principles Improve Prompt Learning In Vision-Language Models

  • 基于贝叶斯原则设计新目标函数,平衡模型适应与泛化
  • 通过先验和后验分布控制微调过程,避免过拟合
  • 适合追求稳定性能的下游任务应用

提示学习因其高效性成为视觉语言模型的热门微调方法,仅需少量可学习参数即可显著提升目标任务性能。然而,现有方法易在微调数据上过拟合,导致泛化能力差。为此,本文提出一种基于贝叶斯学习原理的新训练目标函数,通过在输出逻辑值(logits)上建立先验分布——其均值由预训练模型参数化,后验对应微调后的模型——实现适应性与泛化性的平衡。该方法使模型在适应下游任务的同时,仍保持与预训练模型的接近度,有效缓解过拟合问题。

原文摘要 · Abstract (English)

Prompt learning is a popular fine-tuning method for vision-language models due to its efficiency. It requires a small number of additional learnable parameters while significantly enhancing performance on target tasks. However, most existing methods suffer from overfitting to fine-tuning data, yielding poor generalizability. To address this, we propose a new training objective function based on a Bayesian learning principle to balance adaptability and generalizability. We derive a prior over the logits, where the mean function is parameterized by the pre-trained model, while the posterior corresponds to the fine-tuned model. This objective establishes a balance by allowing the fine-tuned model to adapt to downstream tasks while remaining close to the pre-trained model.

提示学习贝叶斯方法视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。