arXiv:2412.10817cs.CVcs.AI2024-12AAAI被引 117

用可控噪声提升视觉语言模型对齐,小样本任务效果显著

Enhance Vision-Language Alignment with Noise

论文配图:Enhance Vision-Language Alignment with Noise
图 1 · 摘自论文原文
  • 通过可学习的有益噪声注入视觉与文本编码器
  • 在11个数据集上小样本分类性能显著优于基线
  • 适合资源受限场景下的视觉语言模型微调

随着预训练视觉-语言(VL)模型的发展,如何在下游任务中增强视觉与语言模态间的对齐成为关键挑战。不同于现有方法通过增加模块来改进双模态表示,本文探索在冻结模型基础上,通过定制化噪声进行微调的可行性。受有益噪声(正激励噪声,Pi-noise)的科学启发,我们提出一种新范式:学习能促进对齐的有益噪声分布。针对基于CLIP的小样本分类任务,重构其推理过程并引入变分推断,实现对视觉与语言模态的π-噪声生成。进而提出正激励噪声注入器(PiNI),通过向视觉和文本编码器注入噪声实现对CLIP的微调。由于可学习有益噪声分布,该方法能在有限算力下获得更丰富的多模态嵌入,从而更好对齐双模态表示。在11个数据集上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

With the advancement of pre-trained vision-language (VL) models, enhancing the alignment between visual and linguistic modalities in downstream tasks has emerged as a critical challenge. Different from existing fine-tuning methods that add extra modules to these two modalities, we investigate whether the frozen model can be fine-tuned by customized noise. Our approach is motivated by the scientific study of beneficial noise, namely Positive-incentive Noise (Pi-noise or $π$-noise) , which quantitatively analyzes the impact of noise. It therefore implies a new scheme to learn beneficial noise distribution that can be employed to fine-tune VL models. Focusing on few-shot classification tasks based on CLIP, we reformulate the inference process of CLIP and apply variational inference, demonstrating how to generate $π$-noise towards visual and linguistic modalities. Then, we propose Positive-incentive Noise Injector (PiNI), which can fine-tune CLIP via injecting noise into both visual and text encoders. Since the proposed method can learn the distribution of beneficial noise, we can obtain more diverse embeddings of vision and language to better align these two modalities for specific downstream tasks within limited computational resources. We evaluate different noise incorporation approaches and network architectures of PiNI. The evaluation across 11 datasets demonstrates its effectiveness.

视觉语言噪声注入小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。